Senior Site Reliability Engineer
Listed on 2026-07-13
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Network Engineer
Are you passionate about cutting edge technology?
Does building next generation Cloud Computing technology excites you?
Join our Compute Site Reliability team!
Our team is responsible for monitoring and measuring the reliability of our suite of Compute products and platform. In collaboration with Engineering and Product teams, we focus on improving the performance and reliability of the products we support.
Partner with the best
You’ll apply statistical data analysis and networking knowledge to diagnose and solve some of the Internet’s most difficult content delivery/cloud challenges. You will work with cross-functional teams to benchmark the performance of our key products and services and influence their evolution.
As a Senior Site Reliability Engineer, you will be responsible for:
- Investigating and troubleshoot networking problems within Linux based networking stack
- Monitoring the functioning and performance of the networking infrastructure via Prometheus metric systems and Grafana dashboards
- Solving complex problems in a timely and accurate manner and avoid recurrence through proactive troubleshooting, automation, and systems programming
- Building software tools and systems to automate analytical tasks and workflows to increase efficiency and reliability.
- Leveraging skills in data analysis, network diagnostics and debugging tools to characterize performance and recommend improvements.
Do what you love
To be successful in this role you will:
- Have 5+ years' experience in Site Reliability or System Engineering role, and bachelor's degree in computer science or related field
- Have expertise in L7 traffic management (Envoy, HAProxy, NGINX) in large-scale distributed systems.
- Be proficient in coding with Python, Perl, R, Java, or SQL & have networking knowledge including routing, firewalls, and DNS
- Have experience with Linux systems and tools such as netstats, trace route, tcpdump
- Be proficient in configuration management and container technologies including Ansible, Salt Stack, Chef, Puppet, Terraform, Docker, Podman, Kubernetes, and Nomad
Benefits at Akamai:
We support your health, well-being, finances, and life beyond work. See our benefits.
Flex Base adapts to your job's needs
Akamai's Flex Base program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).