Senior SRE Engineer
Listed on 2026-08-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Senior SRE Engineer
This is a 6-month contract role, with a potential to convert to full-time on the Platform Engineering team, working closely with global development, operations, and delivery teams. Platform Engineering owns the reliability, availability, and performance of our infrastructure and services across a multi-account, multi-tenant AWS environment spanning multiple regions. The role balances hands-on technical work with team leadership: driving initiatives, mentoring engineers, owning SLA/SLO adherence, leading incident management, and automating toil out of day-to-day operations.
Ready to lead reliability at scale? If you thrive in fast-paced environments and are passionate about system reliability, security, and operational excellence, we invite you to join our Platform Engineering team.
As a Senior SRE Engineer, you will:
- Design, implement, and maintain scalable systems for uptime, resilience, and performance, while defining and enforcing SLOs, SLIs, and SLAs with product teams.
- Own incident detection, escalation, and resolution, developing playbooks for failure scenarios and leading post-mortem analyses with corrective actions.
- Build monitoring, logging, and alerting systems, and develop automation for deployments, maintenance, and toil reduction.
- Drive CI/CD improvements and maintain operational documentation, runbooks, and best practices.
- Partner with developers on reliability by design, fostering shared responsibility for observability and fault tolerance.
- Enforce security policies and standards, collaborating on vulnerability identification and mitigation strategies.
Required Experience & Skills
- AWS Cloud Services:
Deep working knowledge across compute, networking, security, storage, and data services in multi-account, multi-region environments. - Kubernetes/EKS:
Experience deploying and maintaining Kubernetes with multi-AZ, multi-region, and multi-tenant topologies. - Hybrid Networking:
Expertise in VPC design, VPN connectivity, on-prem/hybrid topology, and edge solutions. - CI/CD Pipelines:
Experience building and maintaining pipelines using Code Pipeline, Code Build, ECR, and ArgoCD. - Infrastructure Design:
Ability to design scalable, resilient infrastructure with security, observability, and automation built in. - Cert/PKI Management:
Knowledge of certificate and PKI lifecycle management, including HSM-backed key protection and mTLS. - Security & Cryptography:
Strong understanding of encryption, cryptographic protocols, and HSM integration. - Scripting & Automation:
Proficiency in Python, Go, and Bash for automation and infrastructure-as-code. - Compliance as Code:
Experience using AWS Config to translate policy requirements into automated compliance checks. - Observability Platforms:
Hands-on experience with Datadog for metrics, distributed tracing, log management, and APM. - Chaos Engineering:
Experience applying chaos engineering practices to proactively validate system resilience.
Preferred Skills & Attributes
- Tactical & Operational:
Ability to drive team initiatives, own delivery outcomes, and act as a reliability champion. - System Reliability Advocacy:
Champion robust system design and operational excellence across engineering and product teams. - Analytical Problem-Solving:
Skilled at diagnosing and resolving complex, multi-layer issues in distributed, cloud-native environments. - Cross-Functional Communication:
Effective collaborator across development, security, and operations teams in a global organization. - Incident Response Leadership:
Experience leading swift resolution of high-severity incidents to minimize downtime. - Culture of Ownership:
Fosters accountability and a blameless culture focused on reliability and continuous improvement. - Adaptability:
Comfortable problem-solving creatively in fast-moving environments with shifting priorities.
The Environment You'll Work In
The Platform Environment
- Multi-account AWS architecture:
Per-client account isolation for highest-tier tenants; shared-account segregation for standard tenants. - Multi-region deployment:
Active-active or active-passive topology for disaster recovery and geographic resilience. - Hybrid connectivity:
Client API integration via VPN to on-premises or cloud environments with HSM-backed mTLS. - Service mesh:
Istio on EKS for traffic management, mutual TLS, observability, and policy enforcement. - Git Ops delivery:
ArgoCD for declarative, version-controlled application deployment to Kubernetes. - End-to-end observability:
Datadog across all tiers for real-time visibility, alerting, and capacity planning.
The Technology Stack
- Cloud & Networking:
Route
53, ELB (NLB/ALB), WAF, VPC, NAT, IGW, VGW, VPN, Lambda - Compute & Orchestration: EKS (Kubernetes), Istio Service Mesh, NGINX, ArgoCD
- Data & Messaging: RDS, DynamoDB, S3, ActiveMQ, Redis
- Security & Compliance: AWS HSM, KMS, Secrets Manager, Parameter Store, AWS Config, mTLS/PKI
- CI/CD & Registry:
Code Pipeline, Code Build, ECR - Observability:
Datadog (metrics, logs, traces, APM) - Scripting & IaC:
GoLang, Python, Terraform
Our Commitment
Akoya is…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).