×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, Data & Caching Systems

Job in San Mateo, San Mateo County, California, 94404, USA
Listing for: Ursus
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, IT Support, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below

Site Reliability Engineer - Data & Caching Systems

AI systems require efficient data handling  are seeking an individual passionate about optimizing software, hardware, and data transfer to support our AI initiatives across the country. The IT storage team manages petabytes of on-premises, clustered POSIX storage for AI modeling and is developing the next-generation storage solutions. This includes building a geo-distributed file system/data lake to support autonomous robotaxis operations nationally and globally.

Our initial focus is on a high-performance caching system significantly outperforming AWS S3.

Responsibilities:
  • Onboard new customers onto our caching solution, helping them update their applications and verifying successful integration.
  • Help scale and operate our fleet of caching servers throughout the world by leveraging automation. (Terraform, Spacelift, Salt Stack, Ansible, etc.)
  • Help improve observability and monitoring of the caching system. (OTEL, Grafana, Zabbix, etc.)
  • Improve CI/CD workflows and help with the software release cycle. (Git Hub, Buildkite, Bamboo, Bazel, etc.)
  • Improve user-facing documentation and operational runbooks for the caching system.
  • Report bugs to software developers. (Jira, Confluence, etc.)
  • Optionally help fix bugs in the caching system. (Rust)
Qualifications:
  • 2+ years of experience operating distributed services.
  • Excellent written and verbal communication skills.
  • Able to organize and work on several long-running processes at one time.
  • Ability to troubleshoot complex problems between the system and customer’s application.
  • Expertise in automation and Infrastructure as Code (IoC). (Terraform, Spacelift, Ansible, Salt Stack, etc.)
  • Expertise in monitoring, alerting, and observability platforms. (OTEL, Grafana, Zabbix, Ops Genie, Incident.io, etc.)
  • Motivated to learn new technologies and think differently.
Bonus

Qualifications:
  • Experience with Python, C++, Rust, or similar programming languages.
  • Experience with CI/CD workflows. (Buildkite, Bamboo, Git Hub, etc.)

Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate or annual salary only, unless otherwise stated. In addition to base compensation, full-time roles are eligible for Medical, Dental, Vision, Commuter and 401K benefits with company matching. IND
123

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary