Senior/Cloud Reliability Engineer
Job in
Mountain View, Santa Clara County, California, 94039, USA
Listed on 2026-09-04
Listing for:
ThoughtSpot
Full Time
position Listed on 2026-09-04
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
Job Description & How to Apply Below
We are seeking a Staff Site Reliability Engineer with deep enterprise SaaS operations expertise to own the availability, reliability, security, and efficiency of our Multi-Cloud (AWS, GCP) production SaaS platform. The ideal candidate brings hands‑on experience running highly available, large-scale Kubernetes‑based control and data planes, a strong bias toward automation and AI‑augmented operations, and a proven track record in production security, capacity management, and cloud‑native data infrastructure.
Responsibilities- Operate a high-scale, multi-cloud (AWS, GCP) SaaS platform — ensuring reliability, performance, and uptime for business‑critical production workloads.
- Embed AI and Agentic workflows into SRE practice: leverage AI Ops platforms and LLM‑powered autonomous agents for anomaly detection, automated triage, runbook execution, and incident summarization to reduce MTTR.
- Drive capacity planning and scaling operations — proactively model growth, right‑size infrastructure, and implement horizontal/vertical autoscaling strategies to support SaaS growth without reliability regression.
- Architect and operate Kubernetes controller frameworks governing both control plane and data plane services; define and enforce operational standards for cluster lifecycle, workload scheduling, autoscaling, and failover.
- Own operations of high‑scale cloud‑native databases and data infrastructure:
PostgreSQL/RDS, DynamoDB, MySQL, Elasticsearch/Open Search, Elasti Cache (Redis/Memcached) on AWS and GCP — including performance tuning, backup/recovery, and incident response. - Lead incident response and blameless post‑mortems for P0/P1 events; drive root cause analysis to permanent resolution and prevention — eliminating repeat incidents through systemic fixes, not workarounds.
- Define and enforce a culture of automation‑first: identify and eliminate toil through self‑healing systems, automated remediation pipelines, and infrastructure‑as‑code (Terraform, Helm, Git Ops).
- Participate in on‑call rotations for critical cloud infrastructure; serve as a senior escalation point and incident commander during high‑severity events.
- Achieve quantifiable SaaS operational Excellence measured by related SLI/SLO/SLA
- B.Tech. degree in Computer Science or equivalent.
- At least 6+ years of Enterprise SaaS Ops experience
- Strong proficiency in programming, particularly with Go and Python, and experience with Infrastructure as Code (IaC) tools like Terraform and Ansible.
- Expertise in Cloud Security and/or Cloud networking
- Experience with AI Ops tools, Agentic LLM.
- Experience/ Knowledge in Cloud Services, Kubernetes, Cloud Databases like Postgres/RDS/MySQL/DynamoDB, Elastic, Kafka, and Microservice architecture is a bonus.
- Experience in implementing and operating enterprise‑grade observability ( metrics, logs, tracing), alerting stack in a Cloud SaaS environment
- Strong debugging and problem‑solving skills (network, systems, database, and application).
- Advanced professional certifications from Cloud Providers ( AWS, Azure, GCP) in domains like K8s, Solution architecture, networking, and databases are a bonus.
- Full Stack Architecture/Development Experience is a bonus.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×