Cockroach DBA
Listed on 2026-08-08
-
IT/Tech
Disaster Recovery IT, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Design, deploy, administer, and maintain production-grade Cockroach DB clusters across cloud and on-premises environments.
Monitor database health, performance, latency, throughput, and resource utilization to ensure service reliability and availability.
Implement and manage backup, restore, disaster recovery, and business continuity strategies.
Perform database capacity planning, performance tuning, indexing, and query optimization.
Develop automation scripts and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks.
Establish and manage SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
Drive incident management, root cause analysis (RCA), postmortems, and preventive remediation activities.
Build and maintain monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms.
Collaborate with Dev Ops and Engineering teams to improve platform reliability, scalability, security, and operational excellence.
Support production releases, database migrations, version upgrades, and platform modernization initiatives.
Participate in on-call rotation and provide support for critical production incidents.
- Implement database security controls, access governance, auditing, and compliance best practices.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).