Site Reliability Engineer IV
Listed on 2026-10-02
-
Software Development
About Expert Voice
We work with hundreds of the world’s most respected consumer brands – companies like Garmin, Carhartt, Ariat, Taylor Made, Brooks, REI and more – to engage influential everyday experts, and to build, track and reward their helpful expertise.
Job TypeFull-time
DescriptionAbout Expert Voice
We work with hundreds of the world’s most respected consumer brands – companies like Garmin, Carhartt, Ariat, Taylor Made, Brooks, REI and more – to engage influential everyday experts, and to build, track and reward their helpful expertise.
At Expert Voice, we are proud to have cultivated a community of individuals who embody our core values. Expert Voice employees are authentic, driven, bold and give a damn. We believe that our people are the heart and soul of our organization and their commitment to these values are what drives our success. When you join Expert Voice, you become part of a team that is united by a shared passion for excellence and a desire to make an impact.
SummaryExpert Voice is seeking a Site Reliability Engineer IV to set the technical direction for how we run, automate, and scale our platform. This is a senior leadership-track engineering role for someone who can see where the team needs to be and builds the architectural runway to get there. You’ll pair deep hands‑on ability with forward‑looking judgment, owning hard systems end‑to‑end while championing AI and automation to eliminate the manual work that slows engineering teams down.
You come from a background that understands how software is built, not just how it’s operated.
You’ll join a small, senior team where each engineer owns a domain, partnering across infrastructure, the developer platform, and security to raise the bar for the whole team.
Key Responsibilities- Lead incident response for complex outages: dig past surface symptoms to true root cause across logs, metrics, trends, and database behavior, drive resolution under pressure, and turn findings into durable fixes and better runbooks.
- Participate in the on-call rotation: a full week at a time, shared with three other engineers, so on‑call comes around every fourth week.
- Debug and troubleshoot across the stack, including code from other teams, and ensure the team’s work stays documented and maintainable.
- Own CI/CD: keep build and deploy pipelines fast and trustworthy, with deploys and rollbacks that any engineer can run.
- Own systems end‑to‑end (architecture, design, optimization, and long‑term maintenance), accountable for reliability and performance at scale.
- Champion AI and automation across our operations: automate away repetitive operations work and the one‑off manual steps that never got fixed, and build the tooling and prototypes that prove what's possible.
- Lead architecture discussions and design reviews, and drive adoption of new technologies and patterns.
- Partner across infra, dev platform, and security, strengthening those areas rather than working around them.
- Set the technical vision and architectural runway: a forward technology plan (debt and investment) that keeps the team ready to deliver what’s next.
- Mentor engineers and raise the technical bar across the team.
- 10+ Years SRE/Dev Ops/systems experience with a track record of owning and architecting production systems at scale.
- Strong coding ability across the stack (Python, Go, or Java; comfortable in Bash).
- Deep Linux and networking expertise (TCP/IP, DNS, HTTP/TLS) and the ability to design for performance at scale.
- Strong AWS, container (Docker), and infrastructure-as-code (Terraform) experience.
- A track record of setting technical direction and driving an architectural plan to execution.
- Strong critical-thinking and troubleshooting instincts: able to work from symptoms…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).