Site Reliability Engineer
Listed on 2026-08-24
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Site Reliability Engineer
Austin, Texas, United States;
Reston, Virginia, United States
We are looking for a talented Senior Site Reliability Engineer to join our team to deliver world class search technologies to mobile devices. You will be working with a smart team of Engineers to lead and drive the stability, reliability, and observability of all Seekr's software, powering Seekr's powerful search technology.
From your first day, you will make a valuable — and valued — contribution. We are a fast-growing company where no one is a bystander. We offer you the opportunity to delight millions of consumers around the world while gaining meaningful experience across a variety of disciplines.
Duties and Responsibilities- Participate in the design, architecture and implementation of systems, software, networks, and services required to keep the Seekr platform within SLAs.
- Lead development of solutions to complex operational and reliability challenges and proactive detection of system failures and scalability issues in production.
- Work in close collaboration with software development teams to build, monitor and triage our services.
- Load testing applications across the organization.
- Take ownership of our observability stack.
- Share ownership of our incidence response practices and procedures.
- Identify and remedy operational inefficiencies through automation.
- Troubleshoot and assist with incident response on live systems and deployment related issues.
- 5+ years working in Site Reliability Engineering preferably managing SaaS environments.
- 5+ years of experience working with Linux systems, familiar with Linux fundamentals including network protocols.
- Experienced with self-hosted monitoring, metrics, and centralized logging tools (ELK, Prometheus, InfluxDB, Grafana; required).
- Skilled in load testing applications.
- Proficient in programming and/or scripting languages (Python, Ruby, Bash, Java).
- Adept with container technologies Docker and Kubernetes (required).
- Demonstrated ability in systems configuration management using automation tools such as Puppet, Chef, Ansible, and Terraform (Puppet and Terraform preferred).
- Knowledgeable in monitoring Kubernetes, Elasticsearch, Kafka, and Aerospike (preferred).
- Versed in Git (on-prem or Git Hub); familiarity with Git Lab and ArgoCD is a bonus.
- Comfortable working within a hybrid cloud environment, with on-prem experience preferred.
- Proven track record of delivering results on time and with high quality.
- Reliable in maintaining SLAs.
- Effective communicator with the ability to work across multiple business and technical teams.
- Advanced degree is nice to have
Seekr is a leader in explainable and trustworthy artificial intelligence designed to power mission-critical decisions in enterprises, government, and regulated industries. Seekr Flow™, our end-to-end AI platform, provides secure, auditable AI solutions tailored to sectors where transparency, accuracy, and compliance are paramount. Available across cloud, on-premises, and edge environments, Seekr Flow reduces bias, strengthens data integrity, and simplifies model oversight so organizations can rely on trusted AI decisions in high-stakes settings that impact society's most sensitive and vital systems.
Trusted by leading enterprises and government agencies, we partner with defense, finance, telecom, and critical infrastructure leaders to enable AI solutions that drive real-world results with unmatched transparency and control. We are a team of strategic thinkers and problem-solvers tackling the toughest challenges facing critical infrastructure and global enterprises through best-in-class AI models and customer deployment. Our team operates with unwavering commitment to our core values and mission:
- We are driven by outcomes—our customers' success is what we strive for every day.
- We believe trust is earned, which is why we build explainability and transparency into the entire AI lifecycle.
- We take our responsibility to deliver secure AI seriously.
- We believe innovation drives progress—we are building the technologies that power the systems our society depends on.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).