Principal DevOps Architect
Listed on 2026-07-26
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure, Systems Engineer
This role offers the opportunity to define and shape the future of cloud infrastructure for a global technology platform supporting critical healthcare and research solutions. As a senior technical leader, you will architect scalable, secure, and highly reliable cloud environments while driving modern Dev Ops practices across the organization. You will combine deep hands‑on engineering expertise with strategic architectural vision, influencing teams through technical excellence and innovation.
The position focuses on infrastructure automation, AI platform operations, reliability engineering, and compliance‑driven cloud solutions. You will play a key role in advancing infrastructure as code, improving developer experiences, and ensuring systems remain secure, efficient, and audit‑ready. This is an impactful opportunity for an experienced architect who thrives in complex environments and enjoys solving challenging technical problems.
The Principal Dev Ops Architect will serve as the senior technical authority responsible for designing, implementing, and evolving cloud platforms, Dev Ops practices, and AI infrastructure. This hands‑on individual contributor role will influence engineering standards, improve operational excellence, and enable teams to deliver reliable and secure solutions.
- Own the cloud platform architecture, infrastructure roadmap, deployment strategies, observability practices, and AI platform direction while partnering with engineering leadership.
- Establish engineering standards, reference architectures, and best practices for infrastructure as code, CI/CD pipelines, cloud operations, and AI tooling adoption.
- Design, maintain, and optimize cloud infrastructure using Terraform as the source of truth, including reusable modules, automated deployments, and policy‑driven controls.
- Build and operate scalable AWS environments across development, testing, staging, and production while ensuring performance, availability, security, and compliance.
- Develop and improve CI/CD pipelines, release processes, containerized workloads, and deployment automation using technologies such as Docker, Kubernetes, Git Hub Actions, and related tools.
- Lead reliability initiatives by defining SLIs, SLOs, error budgets, monitoring strategies, and incident response processes.
- Operate and govern AI/ML platforms, including model‑serving infrastructure, AI observability, security controls, cost management, and responsible AI practices.
- Implement security and compliance measures supporting healthcare data protection, audit readiness, secrets management, vulnerability management, and supply‑chain security.
- Collaborate with product, engineering, QA, support, and operations teams to improve automation, documentation, and cross‑functional delivery.
- Mentor engineers through technical leadership, architecture reviews, and hands‑on contributions without direct management responsibility.
The ideal candidate is a highly experienced cloud and Dev Ops professional with a strong background in architecture, automation, security, and large‑scale platform engineering. They should be comfortable leading technical initiatives, influencing teams, and delivering solutions in regulated environments.
- Bachelor’s degree in software engineering, computer science, or equivalent technical experience.
- 10+ years of experience in Dev Ops, SRE, platform engineering, cloud architecture, or related disciplines, including senior individual contributor or architect‑level responsibilities.
- Proven experience designing and operating large‑scale cloud environments, preferably on AWS.
- Strong hands‑on expertise with Terraform, infrastructure as code practices, reusable modules, automated provisioning, and CI/CD‑based infrastructure management.
- Experience building and managing CI/CD pipelines using tools such as Git Hub Actions, Jenkins, Git Lab, or AWS-native solutions.
- Strong knowledge of Docker, Kubernetes, container orchestration, Linux administration, scripting, and automation.
- Experience with cloud observability, monitoring, troubleshooting, Open Telemetry, and reliability engineering practices.
- Familiarity with security and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).