Senior Site Reliability Engineer El Segundo, CA; On-site
Listed on 2026-09-20
-
Software Development
AWS
Hive Watch is a tech-forward, inclusive organization fostering the evolution of the physical security industry. We are a diverse team of forward thinkers who empower each other to find creative and collaborative solutions in an industry ripe for modernization. We are passionate about the problems we’re solving for our customers and equally passionate about the company we’re building.
Hive Watch is here to help security teams pivot from chasing threats to preventing them. We protect organizations, people, and property through the intelligent orchestration of physical security programs. With better communication, more insights, and less “noise”, we are modernizing what it means for businesses and their employees to truly feel safe.
POSITION OVERVIEW:Hive Watch is seeking a Senior Site Reliability Engineer to join our Platform Team, where you'll build and operate mission-critical edge infrastructure that connects our SaaS platform to customer systems. You'll help ensure exceptional performance, reliability, and observability across our distributed environment.
WHAT YOU'LL DO:Improve the reliability of mission-critical systems including production monitoring, alerting, and capacity planning
Debug and resolve complex production issues across the full stack, from infrastructure to application code
Participate in a regular on-call rotation to provide 24/7 coverage for critical systems
Perform root cause analysis requiring deep code-level investigation and implement preventive measures
Build automation and tooling to reduce operational toil and improve system reliability
Maintain CI/CD pipelines, observability infrastructure, and database performance optimization
Increase the resiliency, scalability, and maintainability of production environments
Maintain on-call procedures and disaster recovery processes
Contribute to on-call runbooks, postmortems, and reliability best practices
Languages:
Kotlin, Rust, Type Script, and PythonDeployments:
Git Hub Actions, Terraform, Terragrunt, and HelmInfrastructure: AWS (Kinesis, Serverless, RDS, EKS), Kubernetes, Docker, Postgres, IoT Edge, Red Hat Enterprise Linux, Rocky Linux
Strong experience with AWS architecture and services
Experience in physical security, IoT, or edge computing environments
Experience with advanced AWS services (Kinesis, Lambda, EKS, RDS)
Experience with Terraform and Terragrunt specifically
Background in high-availability, multi-tenant SaaS environments
Experience participating in incident response and post-mortem processes
Knowledge of security best practices and compliance requirements
Experience with edge computing and distributed system architectures
Previous experience in a startup or high-growth environment (50-200 employees)
Experience with our tech stack:
Kotlin, Rust, Type Script, Python
5+ years of software engineering experience with strong coding skills in production environments
3+ years of SRE, Dev Ops, or production operations experience
Strong experience with cloud platforms (AWS preferred) and containerized applications (Docker, Kubernetes)
Experience with Infrastructure as Code (Terraform, Cloud Formation, or similar)
Proficiency in at least one object oriented programming language in our tech stack (Java, Kotlin, Python)
Hands-on experience with relational databases and SQL performance optimization
Experience with monitoring and observability tools (Prometheus, Grafana, Data Dog, or equivalent)
Strong debugging skills across distributed systems and microservices architectures
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
Salary range for this position: $160,000 to $200,000 per year
Eligible to participate in Hive Watch Equity…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).