Site Reliability Engineer
Listed on 2026-08-28
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Site Reliability Engineer (SRE)
We are seeking an experienced Site Reliability Engineer (SRE) to join our Product Platform team. This role will serve as a critical bridge between product development teams and the platform engineering organization, helping translate application infrastructure needs into scalable, reliable platform solutions.
The ideal candidate has deep, hands-on experience with Kubernetes, AWS, and observability, along with strong troubleshooting and cross-functional communication skills. This is an infrastructure-focused engineering role for someone who is comfortable working directly with product teams, diagnosing complex application and platform issues, and driving problems through resolution across multiple technical teams.
Responsibilities- Build, manage, maintain, and troubleshoot Kubernetes clusters and containerized application environments.
- Partner closely with product and application development teams to understand infrastructure requirements and translate them into actionable platform engineering needs.
- Serve as a first point of contact for infrastructure and reliability issues affecting product teams, performing root-cause analysis across application, Kubernetes, cloud, networking, and security layers.
- Resolve Kubernetes and platform-related issues directly while coordinating with other infrastructure, cloud, network, and security teams when issues fall outside the platform team's ownership.
- Create, configure, and maintain AWS resources supporting application and platform environments.
- Support and enhance an internal observability platform used to monitor applications, services, and infrastructure.
- Onboard new applications and use cases into the observability platform by partnering with technical teams to understand monitoring and telemetry requirements.
- Develop and maintain Python scripts used for infrastructure automation, troubleshooting, platform operations, and observability.
- Support CI/CD and Git Ops-based deployment processes for Kubernetes environments.
- Improve platform reliability, scalability, monitoring, operational efficiency, and developer experience.
- Participate in troubleshooting and root-cause investigations involving multiple engineering teams and drive issues through successful resolution.
- Document platform standards, troubleshooting procedures, operational processes, and technical solutions.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field with 4+ years of relevant experience, or 6+ years of equivalent professional experience in lieu of a degree.
- Deep hands-on experience with Kubernetes, including building clusters from scratch, cluster administration, deployments, networking, and troubleshooting.
- Strong hands-on experience with AWS, including creating and maintaining cloud infrastructure and resources.
- Experience working with AWS services such as S3 and RDS.
- Strong understanding of observability, monitoring, logging, metrics, and application performance concepts.
- Experience with observability platforms such as Datadog, Splunk, Grafana, Prometheus, or similar technologies.
- Experience using Python for scripting, automation, or infrastructure-related tasks.
- Strong troubleshooting and root-cause analysis skills across complex application and infrastructure environments.
- Excellent verbal and written communication skills with the ability to work effectively across product, application, infrastructure, cloud, networking, and security teams.
- Ability and willingness to work in a highly collaborative position that combines hands-on engineering with significant cross-team coordination.
- Experience with Helm for Kubernetes application packaging and deployment.
- Experience with Argo CD and Git Ops-based deployment practices.
- Experience building or maintaining CI/CD pipelines.
- Experience with Prometheus, Grafana, and Open Telemetry.
- Experience with Amazon EKS and IAM.
- Familiarity with Istio or other service mesh technologies.
- Experience with Kafka or Amazon MSK.
- Experience with infrastructure-as-code and configuration-management technologies such as Terraform or Ansible.
- Experience supporting internal developer platforms, infrastructure platforms, or observability products.
- Experience onboarding development teams or applications onto centralized platform services.
- Experience working in regulated or highly controlled technology environments.
- Interest in learning and taking ownership of application code supporting internal platform tooling.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).