×
Register Here to Apply for Jobs or Post Jobs. X

Google Cloud Site Reliability Engineer; GCP SRE

Job in Buffalo Grove, Lake County, Illinois, 60089, USA
Listing for: Iconma
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Position: Google Cloud Site Reliability Engineer (GCP SRE)

Google Cloud Site Reliability Engineer (GCP SRE)

Our client, a IT Services and Consulting company, is looking for a Google Cloud Site Reliability Engineer (GCP SRE) for their Buffalo Grove, IL/Hybrid location.

Responsibilities:
  • Responsible for incident detection & logging and meeting agreed SLA for incident tickets.
  • Responsible for bridge activation & communication (P1–P2).
  • Postmortem preparation (within 24–72 hours) & root cause analysis.
  • Responsible for critical monitoring activities, problem management & Grafana integration.
  • Participate in oncall rotations, handle incidents, and drive timely mitigation and recovery.
  • Automating operational work so services can scale without manual toil also operating highly available, low latency & secure systems.
  • Defining and measuring reliability through SLIs/SLOs and error budgets.
  • Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services.
  • Tune alerting to reduce noise while ensuring rapid detection of user impacting issues.
  • Lead or contribute to post incident reviews and root cause analysis and ensure follow up actions are implemented to prevent recurrence.
  • Added advantage if resource is familiar on tools Tidal, Service Now, Xmatters, Abinitio, Tableau, Opsgenie & Zeke.
Requirements:
  • 8+ years in site reliability engineering, Dev Ops, or cloud engineering.
  • Strong experience supporting production GCP environments.
  • Experience with enterprise incident management and cloud operations.
  • Knowledge/experience in GCP (Big Query, Cloud storage, Dataproc, GKE, Airflow/Composer, Pub-sub, Cloud function, Cloud SQL etc).
  • Knowledge/experience in Github & Visual Studio code.
  • Knowledge/experience in MS-Copilot.
  • Knowledge/experience in Prometheus, Grafana & Splunk.
  • Knowledge in Python/Pyspark/Machine learning is an added advantage.
  • Clear written and verbal communication, particularly under pressure (e.g., during incidents).
  • Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority.
  • Strong communication, analytical, knowledge on entire incident management life cycle process, Agile model experience, and problem-solving skills.
  • EDP-SRE (Google Site Reliability Engineer).
  • Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services.
  • Years of

    Experience:

    10.00 Years of Experience.
Why Should You Apply?
  • Health benefits.
  • Referral program.
  • Excellent growth and advancement opportunities.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary