×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer, GovCloud

Job in McLean, Fairfax County, Virginia, USA
Listing for: Medallia
Part Time position
Listed on 2026-08-15
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
Job Description & How to Apply Below

Senior Site Reliability Engineer

Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.

We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.

We empower exceptional people to create extraordinary experiences together.

Bring your whole self.

The Role and Team

We are growing our Gov Cloud team and looking for a Senior Site Reliability Engineer to help operate and improve Medallia's US public-sector cloud platform. You will support federal agencies and other regulated customers in a highly available, secure, and compliant environment built on AWS Gov Cloud and Kubernetes.

This is a hybrid role based near Tysons, Virginia, with regular in-office collaboration and remote flexibility. You will be a hands-on engineer who owns production systems, partners closely with engineering and security teams, and helps keep the platform reliable as we scale.

Responsibilities
  • Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services.
  • Design and operate AWS cloud networking end-to-end — VPC architecture, subnetting and routing, security groups/NACLs, VPC endpoints/Private Link, Transit Gateway, load balancing, and DNS — for secure, segmented, highly available connectivity.
  • Operate and tune production PostgreSQL — high availability and replication, backups and recovery, query and performance optimization, version upgrades, and capacity planning — as part of the platform's data tier.
  • Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues.
  • Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and Git Ops practices.
  • Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks.
  • Work with software engineering, security, and release management teams to deploy changes safely and resolve production issues.
  • Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
  • Participate in an on-call rotation for production support.
  • Document systems and operational procedures clearly.
  • Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.

Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Qualifications

Minimum Qualifications

  • Must reside in the United States and be legally authorized to work in the US without sponsorship.
  • Bachelor's degree or equivalent experience in Computer Science or a related field.
  • 5+ years of experience in Site Reliability Engineering, platform engineering, Dev Ops, or related production infrastructure roles.
  • Production experience with:
    • Kubernetes
    • Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (GCP), Azure, or a similar public cloud platform.
    • Terraform or comparable infrastructure-as-code tools
    • Git and CI/CD pipelines
    • Linux and foundational systems concepts (networking, DNS, TLS/certificates)
    • PostgreSQL (or comparable relational databases) in production — replication, backups, and performance tuning
  • Programming and Automation:
    Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services.
  • Incident & Change Management:
    Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes.
  • Experience participating in a production on-call rotation.
  • Experience troubleshooting complex technical…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary