×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Naperville, DuPage County, Illinois, 60540, USA
Listing for: ampliFI Loyalty Solutions
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Project Manager
Job Description & How to Apply Below

Site Reliability Engineer (SRE) – Release & Operations

The Site Reliability Engineer (SRE) – Release & Operations is a hybrid technical and process-oriented role responsible for bridging the gap between software development and stable IT operations. This role focuses on owning and optimizing the release lifecycle, driving governance through the Change Advisory Board (CAB), and facilitating cross-team deployment communication. In addition to release engineering, this position requires a hands-on technical expert who will develop automated tooling to reduce operational toil, actively participate in production support and on-call rotations, and ensure the high availability, security, and scalability of AWS-based cloud infrastructure in a compliant corporate environment.

Responsibilities

Release Management & Governance (CAB)

  • Release Lifecycle Ownership:
    Lead and coordinate the end-to-end release lifecycle, including planning, scheduling, staging, deploying, and post-release validation.
  • Change Advisory Board (CAB):
    Act as a key representative and technical coordinator on the Change Advisory Board, defending upcoming releases, evaluating architectural risk, and ensuring all compliance requirements are met prior to production deployment.
  • Cross-Functional Communication:
    Serve as the primary point of contact for developers, QA, product management, and business stakeholders regarding deployment windows, release status, risk assessments, and rollback plans.
  • Process Optimization:
    Standardize and mature release processes, transitioning manual gatekeeping into automated CI/CD guardrails and repeatable workflows.

Site Reliability & Production Support

  • Production Support & On-Call:
    Provide hands-on tier-2 production support, ensuring operational stability and participating in the active engineering on-call rotation.
  • Tooling & Automation Development:
    Design, build, and maintain internal scripts, custom tooling, and automated pipelines (Python, Bash, Power Shell) to reduce operational "toil" and streamline release operations.
  • Incident Response & RCA:
    Respond promptly to system outages, service interruptions, and security alerts. Lead post-incident Root Cause Analysis (RCA) efforts and implement permanent preventive engineering solutions.
  • Infrastructure-as-Code (IaC):
    Maintain, deploy, and scale AWS cloud infrastructure using Terraform, AWS Cloud Formation, or Ansible to ensure environment parity and drift-free deployments.
  • Observability & Monitoring:
    Configure, tune, and optimize monitoring and alerting systems (e.g., Cloud Watch, Prometheus, Grafana) to provide comprehensive visibility into release health and production performance.

Security, Compliance, & Collaboration

  • Compliance Alignment:
    Ensure all deployment and release activities strictly adhere to PCI, SOC2, and internal corporate security/governance standards.
  • Collaboration:

    Work closely with software engineering, QA, and platform infrastructure teams to build "paved paths" for developers to ship software safely and rapidly.
  • Documentation:
    Maintain pristine, audit-ready documentation for release logs, standard operating procedures (SOPs), runbooks, and incident timelines.
Essential Skills and Experience

Experience:

3–5 years of experience in SRE, Dev Ops, system administration, or Release/Operations engineering roles.

Release

Experience:

Proven experience coordinating software releases, managing multi-tier deployment pipelines, and working within formal ITIL/Change Management frameworks (including active CAB participation).

Education:

Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience).

AWS Expertise:
Hands-on experience configuring, deploying, and maintaining AWS infrastructure and serverless architectures (EC2, S3, RDS, IAM, Lambda, API Gateway).

Infrastructure-as-Code (IaC):
Solid understanding and working experience with IaC tools such as Terraform or AWS Cloud Formation.

Scripting & Automation:
Strong proficiency in scripting languages (especially Python and Bash) to write automation tooling and integrate systems.

Incident Management:
Demonstrated experience leading Root Cause…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary