Sr Site Reliability Engineer; Relic/Octopus Deploy/Terraform
Job in
Chicago, Cook County, Illinois, 60290, USA
Listed on 2026-07-27
Listing for:
The Judge Group
Full Time
position Listed on 2026-07-27
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Ready to Lead Observability & Reliability Architecture? At-a-Glance Snapshot
- Role: Senior Site Reliability Engineer / Observability Specialist
- Location / Work Model: Hybrid / Remote (US)
- Compensation / Perks: Competitive Salary + Equity + Comprehensive Benefits
- Massive Financial Scale: Build and optimize high-throughput liquidity and corporate financial platforms handling real-time, global payment flows.
- Greenfield Ownership: Establish core incident management and reliability frameworks from scratch—you'll set the playbook, not just follow one.
- Engineering-First Consulting: Enjoy the balance of hands-on technical execution (80%) and high-impact cross-functional team coaching (20%).
- Engineered Observability: Build and optimize New Relic instrumentation (NRQL, APM, Logs) across Azure and AWS to streamline metrics collection using RED/USE frameworks.
- Infrastructure as Code (IaC): Author and maintain Terraform modules to manage monitoring configurations, alert pipelines, and cloud resources across multi-cloud environments.
- Incident Architecture: Establish early-stage incident response foundations, deploying Incident.
IO, Slack/Ops Genie integrations, and automated escalation workflows to drive down MTTD/MTTR. - Reliability Strategy: Define and govern actionable SLOs/SLIs and error budgets, optimizing log ingestion costs while eliminating alert fatigue for stream-aligned engineering teams.
- Technical Enablement: Partner directly with platform and product teams through hands-on coaching, post-incident reviews (PIRs), and Azure Dev Ops CI/CD pipeline automation.
- 7+ Years in SRE/Dev Ops: Deep experience scaling enterprise cloud infrastructure and production reliability in high-availability environments.
- Observability Mastery: Expert-level hands-on skills with New Relic (NRQL, Synthetics, APM) and defining actionable SLOs/SLIs.
* MUST HAVE* - Strong IaC & Automation: Advanced proficiency in Terraform and Power Shell scripting within enterprise Windows (80%) and Linux (20%) environments.
* MUST HAVE* - Cloud & CI/CD Expertise: Proven track record in Azure (App Services, Virtual Machines, Azure SQL) combined with Azure Dev Ops or Octopus Deploy pipelines.
- Incident Management Focus: Experience designing on-call rotations, runbooks, and incident response operations via Incident.
IO, Pager Duty, or Ops Genie.
Ready to reshape global financial infrastructure and take complete ownership of production reliability?
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×