System Engineer – Site Reliability Engineering; SRE
Listed on 2026-09-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability
Systems Engineer
- Site Reliability Engineering (SRE)
The Systems Engineer
- Site Reliability Engineering (SRE) is responsible for the reliability, scalability, and performance of mission-critical cloud and on-prem services that support millions of Marriott customers globally. This role involves overseeing incident management, driving automation efforts, and working closely with cross-functional teams to ensure alignment between SRE strategy and business objectives. Partners closely with Product Teams, Applications teams, Infrastructure, and the broader Applications and Infrastructure Delivery teams to develop key metrics and KPIs to improve applications stability, availability and performance.
The ideal candidate will bring strong communication skills, collaborating with key stakeholders across the company to optimize cloud infrastructure and uphold the highest standards of operational excellence in a dynamic, fast-paced environment.
Required Education and Experience
- Undergraduate degree in an engineering or computer science discipline and/or equivalent experience/certification
- 5+ years of experience as a Site Reliability Engineer (SRE), building and managing highly available and mission critical systems
- Expertise in AWS services including designing highly available, multi-AZ and multi-region architectures including:
- Compute: EC2, Auto Scaling, Lambda
- Containers: EKS (Mandatory), ECS (good to have)
- Networking: VPC, subnets, route tables, NAT gateways, Transit Gateway
- Security: IAM roles/Policies, KMS, Secret manager
- Storage and Databases: S3, EBS, EFS, RDS, Document DB.
- Deep understanding of SRE practices such as Service Level Objectives, Error Budgets, Toil Management, Observability & Monitoring, Blameless Postmortems, Incident Response Process, Capacity Planning
- Strong understanding of Cloud Security best practices and responsibility model.
- Experience driving cloud cost optimization initiatives (rightsizing, reserved instances, autoscaling strategies, cost observability)
- Proven automation and programming experience in one or more of the following languages:
Python, Bash, Power Shell - Strong working knowledge of modern, continuous development techniques and pipelines (Agile, Kanban, Jira, CI/CD, Helm, Harness, Jenkins, Git, Artifactory, Vault)
- Production level expertise with containerization orchestration engines such as Kubernetes (EKS, AKS, ACK)
- Hands-on experience with service mesh technologies to enable secure and resilient service communication, including mTLS, traffic shaping, and policy enforcement.
- Strong experience troubleshooting API-related issues in distributed systems, including latency, authentication/authorization failures, rate limiting, and upstream/downstream dependency failures.
- Ability to analyze API traffic and debug issues using logs, traces, and metrics.
- Experience with Infrastructure as Code (Iac) tools like Terraform and Cloud Formation.
- Experience with configuration management and automation tools such as Ansible.
- Deep expertise and hands-on experience with Linux administration (RHEL, Ubuntu, CentOS, AWS Linux)
- Solid understanding of Virtualization Technologies (VMware vSphere, KVM etc)
- Strong understanding of networking fundamentals such as Load Balancing, Firewalls, Security Groups, NACLs, TCP/IP, DNS, HTTP/HTTPS, SSL/TLS etc
- Deep understanding and/or experience with Cloud Native, Relational and No
SQL databases like RDS, MySQL, PostgreSQL, Cassandra or Couchbase - Strong experience designing and implementing end-to-end observability solutions across metrics, logs, and traces using tools like Prometheus, Grafana, ELK Stack, and Open Telemetry.
- Proven ability to define SLIs/SLOs, build actionable alerting systems, and leverage telemetry data for incident response, root cause analysis, and performance optimization.
- Experience with deploying, monitoring, and troubleshooting large-scale, distributed applications in cloud environments such as AWS
- Experience in vulnerability management, OS hardening, patching, security compliance of infrastructure, applications and databases
- Experience in implementing OS and cloud hardening guidelines and perform regular vulnerability remediation.
- Famil…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).