SRE Engineer
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
SRE Engineer Onsite /New Jersey Full Time
We are seeking an SRE Engineer focused on observability, Kubernetes, and cloud infrastructure to support our large-scale GCP/AWS/EKS platform. This role is central to improving SLO reliability, logging pipelines, distributed tracing, dashboards, and automated diagnostics across 10,000+ applications running in EKS. Responsibilities include owning the observability stack:
Prometheus/Grafana, Open Telemetry, Loki/ELK/Splunk, Jaeger, Alert manager, SLO frameworks. Build intelligent monitoring pipelines and ensure high reliability of metric ingestion, log ingestion, tracing, and analytics systems. Develop Terraform modules for observability infrastructure, K8s components, cluster add-ons, and monitoring services. Improve reliability of AWS/GCP/EKS clusters through automation, performance tuning, capacity modeling, and event-driven remediation. Build AI-assisted diagnostics for anomaly detection, auto-alert tuning, automated playbooks, and noise reduction.
Partner with Platform Engineering to ensure Istio/service mesh telemetry, API server health, and node-level insights. Lead operational readiness, SLO reporting, incident management, and root cause analysis for platform outages.
Qualifications include 4-8 years in SRE, infrastructure, or Kubernetes operations. Strong knowledge of EKS/ECS/GKE, Kubernetes internals, and cluster operations. Expertise in observability stacks (Prometheus, OTel, Grafana, ELK, Datadog, Splunk). Advanced Terraform IaC and automation skills (Python/Go preferred). Experience with CI/CD, cloud networking, service mesh (Istio), and capacity planning.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).