More jobs:
Cockroach DB Senior Engineer
Job in
Sunnyvale, Santa Clara County, California, 94087, USA
Listed on 2026-09-17
Listing for:
REALIGN LLC
Full Time, Seasonal/Temporary
position Listed on 2026-09-17
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Disaster Recovery IT, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Job Type: Full Time Job Category: IT Job Description
Role:
Cockroach DB Senior Engineer
Location:
Sunnyvale, CA/ Austin,TX
Fulltime – Permanent
Job Description
Key Responsibilities
- Design, deploy, operate, and scale multi-region Cockroach DB clusters in production environments
- Ensure high availability, fault tolerance, and data consistency for globally distributed clusters
- Monitor cluster health, latency, replication status, and resource utilization using observability tools
- Perform capacity planning and proactive scaling for future growth
- Troubleshoot complex database and infrastructure issues including:
- Network partitions
- Leaseholder and range imbalance
- High latency / throughput bottlenecks
- Implement and test backup, restore, and point-in-time recovery processes
- Automate provisioning, scaling, patching, and upgrades of CRDB clusters
- Perform rolling upgrades with zero or near-zero downtime
- Optimize SQL query performance and database schema efficiency
- Create operational runbooks, SOPs, and on-call playbooks for CRDB
- Participate in on-call rotations and incident response for production clusters
- Design, deploy, operate, and scale multi-region Cockroach DB clusters in production environments
- Ensure high availability, fault tolerance, and data consistency for globally distributed clusters
- Monitor cluster health, latency, replication status, and resource utilization using observability tools
- Perform capacity planning and proactive scaling for future growth
- Troubleshoot complex database and infrastructure issues including:
- Node failures
- Network partitions
- Leaseholder and range imbalance
- Replication lag
- Hotspotting
- High latency / throughput bottlenecks
- Design disaster recovery strategies (multi-region, backup/restore, failover/fallback)
- Implement and test backup, restore, and point-in-time recovery processes
- Automate provisioning, scaling, patching, and upgrades of CRDB clusters
- Perform rolling upgrades with zero or near-zero downtime
- Optimize SQL query performance and database schema efficiency
- Create operational runbooks, SOPs, and on-call playbooks for CRDB
- Participate in on-call rotations and incident response for production clusters
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×