×
Register Here to Apply for Jobs or Post Jobs. X

Observability Engineer​/Site Reliability Engineer

Job in Chicago, Cook County, Illinois, 60290, USA
Listing for: Role, Inc.
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Unix/Linux
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below
Position: Observability Engineer / Site Reliability Engineer

About Ontrac Solutions

Ontrac Solutions is a leading technology consulting firm specializing in solutions that drive business transformation. We partner with organizations to modernize infrastructure, streamline processes, and deliver tangible results.

Role Overview

We are seeking an experienced Observability / Site Reliability Engineer (SRE) to design, scale, and maintain our enterprise monitoring and alerting ecosystems. You will bridge development and operations by ensuring high availability, performance tuning, and deep visibility across distributed multi-cloud and native systems. You will automate infrastructure and build robust observability pipelines using cloud-native tools.

Key Responsibilities
  • GCP & Cloud Management: Architect, optimize, and maintain observability frameworks across cloud environments, focusing on Google Cloud Platform (GCP) observability tools such as Cloud Logging, Cloud Monitoring, Trace, and Profiler.
  • Platform Management: Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Automation & IaC: Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • CI/CD Integration: Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / Open Shift environments using Git Hub, Harness, and other CI/CD pipelines.
  • System Performance: Analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex distributed environments.
  • SRE Evangelism: Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.
Required

Skills & Qualifications
  • Cloud Infrastructure: Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and computing resources.
  • Observability Tooling: Hands-on experience with Grafana, Prometheus, and Google Cloud Observability suites. Knowledge of Grafana Enterprise Metrics (GEM) is desirable.
  • OS & Scripting: Expert-level knowledge of Linux/Unix operating systems with strong shell scripting for automation.
  • Programming: Proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Containers & Orchestration: Experience managing containerized applications on Kubernetes, GKE, and/or Red Hat Open Shift.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary