Senior Cloud/Platform Operations Engineer
Listed on 2026-07-23
-
IT/Tech
SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations
Senior Cloud / Platform Operations Engineer
Location: Boston, MA
We are seeking a Senior Cloud / Platform Operations Engineer to play a key role in the day‑to‑day operations of our AWS and Kubernetes platforms. This is a hands‑on technical role focused on ensuring the reliability, security, and operational excellence of our cloud infrastructure. Unlike traditional platform engineering roles centered on building internal developer platforms or delivering new features, this position is focused on operations, service delivery, and execution.
You'll be responsible for maintaining production Kubernetes clusters, enforcing AWS governance, responding to operational issues, and supporting a high volume of internal requests, as well as participating in an on‑call off‑hours rotation. The ideal candidate thrives in production environments, enjoys solving complex infrastructure problems, and is passionate about building stable, secure, and well‑governed cloud platforms.
- Cloud & Platform Operations
Own the day‑to‑day operations of AWS and Kubernetes environments, ensuring high availability, reliability, and performance. Perform hands‑on administration of production Kubernetes clusters and AWS infrastructure. Troubleshoot complex infrastructure, networking, and container platform issues. Execute platform maintenance activities including upgrades, patching, scaling, and lifecycle management. Continuously improve operational processes, automation, and platform stability.
- Kubernetes Operations
Operate and maintain production Kubernetes clusters. Perform cluster upgrades, patching, and version management. Troubleshoot Kubernetes control plane, worker nodes, networking, storage, ingress, and workload issues. Optimize cluster performance, resource utilization, and resilience. Support containerized application deployments and resolve runtime issues. Implement Kubernetes operational best practices around security, reliability, and scalability.
- AWS Operations & Governance
Support AWS account administration, IAM, networking, and infrastructure management. Implement and maintain AWS governance guardrails, security controls, and access policies. Ensure compliance with organizational standards for cloud infrastructure. Assist with cloud cost optimization and resource management. Partner with security teams to remediate infrastructure risks and vulnerabilities.
- Incident Response & Operational Support
Participate in production incident response and on‑call rotations. Troubleshoot and resolve complex production issues across AWS and Kubernetes environments. Perform root cause analysis and implement corrective actions. Support a ticket‑driven operational model with a strong focus on responsiveness and customer service. Document operational procedures and contribute to knowledge sharing across the team.
- Monitoring & Reliability
Maintain monitoring, logging, and alerting for cloud infrastructure and Kubernetes platforms. Proactively identify reliability risks before they impact users. Improve observability and operational visibility across the platform. Drive continuous improvements to platform uptime and operational efficiency.
- Cross‑Functional Collaboration
Partner with engineering teams to troubleshoot application deployment issues. Support internal users by resolving infrastructure requests and platform‑related issues. Collaborate with security, networking, and development teams to deliver reliable cloud services. Contribute technical expertise during infrastructure planning and operational improvements.
5+ years of experience…
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).