×
Register Here to Apply for Jobs or Post Jobs. X

Senior Systems Reliability Engineer

Job in Mountain View, Santa Clara County, California, 94039, USA
Listing for: ThoughtSpot
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 140000 - 170000 USD Yearly USD 140000.00 170000.00 YEAR
Job Description & How to Apply Below

Systems Reliability Engineer (Technical Support):

About Us:

Thought Spot is an AI-powered analytics platform that enables users to explore and analyze data through natural language queries, making insights accessible to all. Our mission is to deliver reliable, high-performing applications that empower our customers.

The Role:

As part of the Thought Spot SRE team, you will be on the cutting edge of operational intelligence. You will not only ensure service reliability but also act as a trusted partner for our customers — proactively leveraging AI/ML to deliver timely updates, meaningful solutions, and predictive improvements. You are the bridge between our customers and engineering, combining deep systems expertise with a genuine passion for customer success.

If you thrive in dynamic environments and are committed to building resilient, self-optimizing systems, this role is for you.

What You'll Do:
Technical & Customer Support:
  • Act as the primary point of contact for customer-facing technical issues related to our SaaS platform, including data connectivity, report errors, performance concerns, access problems, data inconsistencies, software bugs, and integration challenges.
  • Understand and empathize with the challenges Thought Spot users face, offering tailored solutions to improve their experience.
  • Provide timely, accurate, and clear updates to customers, consistently meeting SLAs and driving issues through to full resolution via tickets and calls.
  • Translate complex technical issues into clear, concise updates for both technical and non-technical stakeholders.
  • Create and maintain knowledge‑base articles to empower customer self‑service and improve support efficiency.
System Reliability & Monitoring:
  • Maintain, monitor, and troubleshoot Thought Spot cloud infrastructure using tools like Grafana, Prometheus, Datadog, and Splunk.
  • Monitor system health and performance through metrics, logs, and dashboards to detect and prevent issues proactively.
  • Implement and leverage AI/ML‑driven solutions for proactive observability, predictive anomaly detection, and intelligent alerting to enhance service reliability and reduce Mean Time to Resolution (MTTR).
  • Understand and apply Net Ops and Sec Ops principles for cloud and on‑premise deployments.
  • Develop and implement automation and best practices to streamline operations and strengthen system reliability.
  • Optimize SRE workflows with AI tools to boost operational effectiveness.
  • Participate in on‑call rotations, lead incident reviews, and conduct thorough root cause analyses to drive continuous improvement.
  • Work cross‑functionally with Engineering to define and implement tools that enhance debuggability, supportability, availability, scalability, and performance.
  • Be an expert in both cloud and on‑premise infrastructure by developing automation and best practices.
What You'll Bring
  • B.S. in Computer Science or equivalent relevant experience.
  • Proven experience troubleshooting complex Linux systems and managing virtualization and cloud platforms (VMware, AWS, Azure, GCP).
  • Hands‑on experience with monitoring tools such as Grafana, Prometheus, Datadog, or Splunk.
  • Demonstrated experience and a keen interest in leveraging AI/ML principles to address SRE challenges — including AIOps, predictive maintenance, and intelligent automation.
  • Prior experience in enterprise customer support, including on‑call rotations and incident management, with the ability to lead root cause analyses.
  • Strong problem‑solving and algorithmic thinking with a solid understanding of system internals.
  • Excellent verbal and written communication skills with the ability to work independently and cross‑functionally in fast‑paced environments.
  • Familiarity with scripting and programming languages such as Python, Go, Bash, or Java.
  • Exposure to infrastructure and service monitoring frameworks with the ability to analyze data to ensure high availability.
Good to Have
  • Experience partnering with Engineering to design and implement mission‑critical tooling and automation that advances system debuggability, high availability, elastic scalability, and performance.
  • Experience with alerting strategies and monitoring system tuning to minimize…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary