Staff Site Reliability Engineer, GovCloud
Listed on 2026-08-15
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
Staff Site Reliability Engineer
Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.
We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.
We empower exceptional people to create extraordinary experiences together.
Bring your whole self.
The Role and TeamWe are growing our Gov Cloud team and looking for a Staff Site Reliability Engineer to help scale how we operate Medallia's US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS Gov Cloud and Kubernetes.
This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. As a Staff Engineer, you will provide technical leadership — improving how we run the platform, mentoring teammates, and driving reliability improvements across teams — while still being hands-on in production.
Success in this role means delivering change safely in a regulated environment: strong automation, clear documentation, disciplined operations, and calm incident response under pressure.
Responsibilities- Design, build, and operate highly available, secure cloud infrastructure on AWS, including networking, identity/access management, Kubernetes clusters, DNS, certificates, and shared platform services.
- Design and operate AWS cloud networking end-to-end — VPC architecture, subnetting and routing, security groups/NACLs, VPC endpoints/Private Link, Transit Gateway, load balancing, and DNS — for secure, segmented, highly available connectivity.
- Operate and tune production PostgreSQL — high availability and replication, backups and recovery, query and performance optimization, version upgrades, and capacity planning — as part of the platform's data tier.
- Ensure the reliability and availability of Medallia applications and infrastructure by monitoring systems, responding to incidents, and eliminating recurring operational problems.
- Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and modern Git Ops practices.
- Improve observability across metrics, logs, and uptime monitoring; tune alerting to reduce noise and speed up diagnosis.
- Partner with software engineering, security, release management, and customer-facing teams to deploy changes safely, resolve production issues, and improve operability.
- Lead or contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
- Participate in an on-call rotation and help improve incident response, communication, and post-incident follow-through.
- Document systems and operational procedures clearly so others can run and improve the platform.
- Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.
- Mentor engineers and help raise engineering standards as the Gov Cloud platform and team grow.
Minimum Qualifications
- Must reside in the United States and be legally authorized to work in the US without sponsorship.
- Bachelor's degree or equivalent experience in Computer Science or a related field.
- 8+ years of experience in Site Reliability Engineering, platform engineering, Dev Ops, or related production infrastructure roles — or 5+ years with demonstrated Staff-level scope (technical leadership, cross-team delivery, incident ownership, and platform/IaC ownership).
- Production experience with:
- Kubernetes
- AWS core services (IAM, compute, object storage, encryption/key management) and AWS cloud networking (VPC design, routing, security groups, load balancing, VPC endpoints/Private Link, Transit Gateway, DNS)
- Terraform or comparable…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).