×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer; remote EMEA)

Remote / Online - Candidates ideally in
San Antonio, Bexar County, Texas, 78201, USA
Listing for: Printify
Remote/Work from Home position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer (remote within EMEA)

About the team:

Platform Infrastructure builds, operates, and continuously evolves FYUL's container platform and cloud foundation. We foster a Dev Ops culture through self‑service tooling, enabling product engineering teams to ship reliable, secure, and cost‑efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure‑as‑code, and acts as the go‑to partner for engineering teams on cloud and Dev Ops topics.

About the role:

We're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go‑to person for our most complex infrastructure problems: you architect and drive large‑scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid‑level SREs. You'll split your time between hands‑on platform work – Kubernetes, AWS, GCP, CI/CD, observability – and technical leadership: proposing designs, reviewing others' work, and helping the team make good build‑vs‑buy and cost/reliability trade‑offs.

Your

daily tasks will include:
  • Infrastructure & reliability: Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.

  • Design and operate our Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.

  • Own and evolve core platform services: cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.

  • Automation & infrastructure as code: Drive large‑scale automation projects and set standards for using Terraform / Terragrunt and Git Ops (ArgoCD) across teams.

  • Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.

  • Observability & incident response
    :
    Be the go‑to person for solving complex, cross‑service infrastructure problems.

  • Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.

  • Participate in on‑call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.

  • Security & cost efficiency
    :
    Lead security efforts within the team – IAM, encryption, secure logging – and mentor others on secure infrastructure practices.

  • Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, Fin Ops practices).

  • Collaboration & mentorship: Mentor mid‑level SREs, provide detailed feedback, and support onboarding of new team members.

  • Communicate complex technical concepts clearly to both engineers and non‑technical stakeholders.

  • Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross‑team initiatives.

Your qualifications:

These reflect the technical bar we hold Senior SRE II's to internally, based on our SRE competency framework and current stack.

  • Core technical experience:
    • Solid Linux systems administration background and comfort scripting in Python.
    • Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well‑Architected Framework; experience in multi‑account AWS environments is a strong plus.
    • Hands‑on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).
    • Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi‑environment management;
      Git Ops experience with ArgoCD.
    • Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.
    • CI/CD experience with Jenkins (Jenkins file, shared libraries) and/or Git Hub Actions, and familiarity with deployment strategies such as blue‑green and canary.
    • Experience with the Grafana observability stack (Grafana, Prometheus, Loki, Tempo, Mimir) – metrics design, dashboarding, alerting, log aggregation, and…
  • Position Requirements
    10+ Years work experience
    To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
    (If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
     
     
     
    Search for further Jobs Here:
    (Try combinations for better Results! Or enter less keywords for broader Results)
    Location
    Increase/decrease your Search Radius (miles)
    0
    200
    Filters
    Education Level
    Experience Level (years)
    Posted in last:
    Salary