More jobs:
Engineer - Cluster Management
Job in
Seattle, King County, Washington, 98127, USA
Listed on 2026-08-22
Listing for:
Tata Consultancy Services
Full Time
position Listed on 2026-08-22
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Software Engineer
Job Description & How to Apply Below
Job Description
- Resource Right‑Sizing & Performance Optimization
- Monitoring & Observability (Prometheus/Grafana/Cloud Watch)
- Manage end-to-end EKS cluster operations, including cluster provisioning, node group configuration, autoscaler tuning, and ongoing maintenance, while implementing memory-optimized configurations across managed node groups and Fargate profiles.
- Perform deep memory profiling of containerized workloads running on EKS: analyzing pod-level memory consumption, identifying over-provisioned resource requests/limits, detecting memory leaks, and understanding how memory is utilized across distributed microservices architectures.
- Implement and validate memory optimization strategies at the container and cluster level, including resource request/limit right‑sizing, Vertical Pod Autoscaler (VPA) configuration, node pool instance type optimization, and memory‑efficient scheduling policies.
- Analyze and optimize how memory is allocated and consumed across distributed systems running on Kubernetes, including container runtime overhead, kernel memory accounting, cgroup memory limits, shared memory segments, and inter-pod communication overhead, to identify hidden memory waste at the infrastructure layer.
- Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.
- Design and execute testing strategies for EKS optimizations, including load testing, soak testing, and chaos engineering experiments to validate that memory-optimized configurations maintain reliability and performance under production‑like conditions.
- Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while communicating effectively across engineering leadership and technical stakeholders.
- Develop comprehensive operational runbooks, SOPs, documentation, and technical specifications that capture EKS optimization patterns and can be consumed by both human engineers and AI agents to orchestrate optimization workflows at scale.
- EKS/Kubernetes Cluster Management
- Memory Profiling & DRAM Optimization
- Resource Right‑Sizing & Performance Optimization
- Linux Container Internals (cgroups/OOM/Runtime Memory)
- Monitoring & Observability (Prometheus/Grafana/Cloud Watch)
- Manage end-to-end EKS cluster operations, including cluster provisioning, node group configuration, autoscaler tuning, and ongoing maintenance, while implementing memory-optimized configurations across managed node groups and Fargate profiles.
- Perform deep memory profiling of containerized workloads running on EKS: analyzing pod-level memory consumption, identifying over-provisioned resource requests/limits, detecting memory leaks, and understanding how memory is utilized across distributed microservices architectures.
- Implement and validate memory optimization strategies at the container and cluster level, including resource request/limit right‑sizing, Vertical Pod Autoscaler (VPA) configuration, node pool instance type optimization, and memory‑efficient scheduling policies.
- Analyze and optimize how memory is allocated and consumed across distributed systems running on Kubernetes, including container runtime overhead, kernel memory accounting, cgroup memory limits, shared memory segments, and inter-pod communication overhead, to identify hidden memory waste at the infrastructure layer.
- Perform custom operations and iterative experiments using internal tooling to validate optimization impact: own end-to-end deployment, test execution, metric validation, and derive actionable insights from results.
- Design and execute testing strategies for EKS optimizations, including load testing, soak testing, and chaos engineering experiments to validate that memory-optimized configurations maintain reliability and performance under production‑like conditions.
- Collaborate with service teams to review cluster architectures, discuss findings, propose optimization plans, and align resolution strategies while…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×