×
Register Here to Apply for Jobs or Post Jobs. X

Pursue Your Passion with Purpose

Job in Raleigh, Wake County, North Carolina, 27601, USA
Listing for: MDA Edge
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, IT Support
Job Description & How to Apply Below

Site Reliability Engineer

This role focuses on reliability, scalability, and performance of enterprise platforms across cloud and on-prem environments. Position requires hands-on engineering with automation, observability, and cross-functional collaboration. Emphasizes metrics-driven improvements and operational excellence under pressure.

Key Responsibilities:

  • Design, implement, and maintain reliable, scalable, secure systems in cloud and on-prem setups.
  • Manage distributed systems on Azure, Linux RHEL7+, and Windows Server 2019+.
  • Build automation workflows using Python, Go, and Bash scripting.
  • Develop Infrastructure-as-Code with Terraform and Ansible.
  • Define, monitor, and refine SLIs, SLOs, and SLAs for service quality.
  • Reduce operational toil through automation and process enhancements.
  • Integrate systems with observability platforms for visibility and proactive issue detection.
  • Troubleshoot incidents, lead response efforts, and conduct post-mortem analyses.
  • Collaborate with software, infrastructure, and business teams for resilient services.
  • Optimize reliability, performance, and maintainability with full ownership.

Required

Skills and Experience:

  • Demonstrate proven experience as Site Reliability Engineer from software engineering, infrastructure, or operations background.
  • Show hands-on expertise with Azure and enterprise OS like Linux RHEL7+ and Windows Server 2019+.
  • Possess strong knowledge of networking and storage including NFS, SAN, and NAS.
  • Understand authentication and naming services such as DNS, LDAP, Kerberos, and Centrify.
  • Exhibit proficiency in Python, Go, Bash scripting, Terraform, and Ansible IaC tools.
  • Design and monitor SLIs/SLOs/SLAs to drive reliability via metrics and automation.
  • Integrate with observability platforms for logs, metrics, and tracing.
  • Remain calm and structured during high-pressure incidents.
  • Display strong communication and collaboration to influence cross-functional stakeholders.
  • Maintain proactive, ownership mindset for continuous improvement.

Key

Skills:

Site Reliability Engineering, Cloud Platforms, Azure, Linux RHEL7+, Windows Server 2019+, Networking Fundamentals, NFS, SAN, NAS, DNS, LDAP, Kerberos, Centrify, Python, Go, Bash, Terraform, Ansible, Infrastructure as Code, Observability Platforms, SLIs, SLOs, SLAs, TOIL Reduction, Incident Response, Post-Mortems, Automation, Metrics-Driven Engineering, System Reliability, Cross-Functional Collaboration, Communication Skills, Ownership Mindset.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary