Site Rel Eng III, GCP
Listed on 2026-08-14
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Network Engineer
We are Optimum, a leader in the fast-paced world of connectivity, and we're seeking driven and enthusiastic professionals to join our team, empower lives, fuel businesses, and drive innovation. Connectivity is no longer a luxury, but a necessity. A career at Optimum means you'll be enabling progress and enhancing lives by providing reliable, high-speed connectivity solutions that keep the world connected.
Our successes, now and in the future, are powered by our amazing product, a commitment to our people and culture, and the connections we make in our communities.
If you are resourceful, collaborative, and passionate about delivering consistent excellence, Optimum is for you!
Job SummaryOptimum is seeking a Lead Site Reliability Engineer (SRE) to design, build, and operate secure, scalable, and highly available platforms on Google Cloud Platform (GCP). This hands-on engineering role combines cloud infrastructure, networking, automation, Kubernetes, and reliability engineering to support critical enterprise and customer-facing applications. The ideal candidate brings deep expertise in GCP, cloud networking, Infrastructure as Code, observability, and incident management, along with a passion for improving reliability through automation and engineering excellence.
You are a hands-on engineer with deep GCP and cloud networking expertise who enjoys solving complex operational challenges through automation, platform engineering, and reliability-focused design. You combine technical depth with leadership to build resilient, scalable cloud platforms that enable teams to deliver software with confidence.
- Design and operate enterprise-scale GCP infrastructure and platforms.
- Build and manage solutions using GKE, Cloud Run, Compute Engine, Cloud SQL, Pub/Sub, Cloud Storage, and Cloud Monitoring.
- Architect and support cloud networking, including Network Connectivity Center (NCC), Interconnect, Cloud VPN, Cloud Router, Shared VPC, and Private Service Connect.
- Develop Infrastructure as Code using Terraform and automate deployment and operational workflows through CI/CD pipelines.
- Define and manage SLOs, improve observability, and reduce operational toil through automation.
- Lead incident
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).