Senior Cloud Platform Engineer
Listed on 2026-09-06
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
Senior Cloud Platform Engineer | Boston, MA | Finance Sector | Hybrid
We are seeking a Senior Cloud Platform Engineer to lead the operational management of our AWS and Kubernetes environments. This role plays a critical part in maintaining platform reliability, strengthening security, and ensuring the seamless operation of production infrastructure.
This is not a traditional platform engineering role focused on building new developer platforms or product features. Instead, the position emphasizes production operations, infrastructure governance, incident management, and operational execution
. The engineer will oversee Kubernetes cluster operations, AWS administration, production support, and incident response while serving as a trusted partner to internal teams and participating in an on-call rotation.
3 days in office (Boston, MA)
Compensation:$150k - $180k + bonus + benefits
Cloud & Platform Operations- Own the day-to-day operations of AWS and Kubernetes environments, ensuring high availability, performance, and reliability.
- Administer and maintain production cloud and container platforms.
- Troubleshoot complex infrastructure, networking, and container-related issues.
- Execute platform lifecycle activities including upgrades, patching, scaling, and maintenance.
- Improve operational processes through automation, standardization, and continuous optimization.
- Operate, maintain, and support production Kubernetes clusters.
- Perform cluster upgrades, patching, and version management.
- Troubleshoot issues across the control plane, worker nodes, networking, storage, ingress, and workloads.
- Optimize cluster performance, resilience, and resource utilization.
- Support containerized application deployments and resolve runtime issues.
- Implement best practices related to Kubernetes security, scalability, and reliability.
- Manage AWS infrastructure, including account administration, IAM, networking, and cloud services.
- Implement and maintain governance controls, security guardrails, and access policies.
- Ensure cloud environments align with organizational compliance and security standards.
- Support cloud cost management and resource optimization initiatives.
Partner with security teams to identify and remediate infrastructure risks and vulnerabilities.
- Participate in on-call rotations and production incident response activities.
- Diagnose and resolve complex issues across AWS and Kubernetes environments.
- Conduct root cause analysis and implement preventative measures.
- Support a high-volume, ticket-driven operational model with a strong customer-service mindset.
- Create and maintain operational documentation, runbooks, and knowledge-sharing resources.
- Maintain and enhance monitoring, logging, alerting, and observability platforms.
- Proactively identify and mitigate reliability risks before they impact users.
- Improve platform visibility and operational insights.
- Drive initiatives that increase uptime, resiliency, and operational efficiency.
In addition to base pay, direct-hire employees may be eligible for client offered benefits such as medical, dental, and vision coverage, and paid leave where required by applicable law. Eligibility may vary based on factors such as location and hire date and is subject to change.
EOE Statement:Specialist Staffing Group is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).