×
Register Here to Apply for Jobs or Post Jobs. X

Senior​/Staff Site Reliability Engineer

Job in Santa Clara, Santa Clara County, California, 95050, USA
Listing for: Gatik AI
Full Time position
Listed on 2026-08-19
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, Network Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 260000 USD Yearly USD 180000.00 260000.00 YEAR
Job Description & How to Apply Below

Senior/Staff Site Reliability Engineer

Santa Clara, CA

About the role

We are seeking an experienced Senior/Staff Site Reliability Engineer to support the operation, monitoring, and scaling of our growing fleet of autonomous vehicles. In this role, you will work closely with our infrastructure and platform teams to manage rollouts of both on-premises and cloud infrastructure in support of expansions to new customer sites. You will be directly involved in the setup and monitoring of our data offload systems, remote supervision stations, and on-prem continuous integration (CI) environments, ensuring our infrastructure is highly reliable, secure, and optimized for performance.

This position plays a critical role in keeping our autonomy operations running smoothly while supporting the rapid growth of our fleet and customer base. This role is onsite 5 days a week at our Santa Clara, CA office!

What you'll do
  • Upgrade and maintain both physical and cloud infrastructure used for offloading data from our autonomous vehicle fleet.
  • Partner with the infrastructure and platform engineering teams to monitor, maintain, and troubleshoot our on-premises data offload and CI systems.
  • Design, develop, and maintain business intelligence (BI) dashboards and ETL (extract, transform, load) pipelines to provide actionable insights into our infrastructure performance and health.
  • Architect and deploy test environments to validate internal and customer-facing infrastructure solutions.
  • Automate deployment, scaling, and upgrading of our remote monitoring software to ensure operational efficiency.
  • Perform ongoing analysis of infrastructure performance, identifying opportunities for optimization in latency, throughput, and reliability.
What we're looking for
  • 5+ years of experience in a related role such as Site Reliability Engineer, Dev Ops Engineer, or Infrastructure Engineer.
  • Strong knowledge of networking fundamentals, including protocols, troubleshooting, and optimization.
  • Hands-on experience with Docker and related ecosystem tools (e.g., Docker Compose, Kaniko).
  • Expertise in Kubernetes deployments and package management via Helm.
  • Proficiency with relational and time-series databases (e.g., Postgres, Timescale DB, InfluxDB).
  • Familiarity with workflow orchestration tools such as Argo and Airflow.
  • Proven experience managing upgrades and rollbacks for customer-facing SaaS environments.
  • Scripting experience in Python and Bash for automation and tooling.
  • Experience building and maintaining dashboards with tools like Grafana.

Salary Range - $180,000- $260,000

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary