Senior Cloud DevOps Engineer
Listed on 2026-08-09
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Infrastructure
xAQUA is building a next-generation, AI-native unified data management platform that brings together enterprise data integration, governance, quality, analytics, semantic intelligence, and generative AI capabilities.
We are looking for a highly experienced Senior Cloud Dev Ops Engineer to design, automate, secure, and operate the cloud infrastructure supporting the xAQUA platform and its customer deployments. This is an opportunity to work on a modern, cloud-native platform serving complex enterprise and government environments.
Role OverviewThe Senior Cloud Dev Ops Engineer will own cloud architecture, infrastructure automation, Kubernetes operations, CI/CD, observability, security, release engineering, and platform reliability across Google Cloud Platform and AWS.
Strong hands-on expertise in both GCP and AWS is required. Experience with Microsoft Azure is highly preferred.
This is a senior individual-contributor role requiring strong technical judgment, hands‑on implementation ability, ownership, analytical problem-solving, and the ability to guide other engineers.
Key Responsibilities:Cloud Architecture and Infrastructure
- - Design, implement, and manage secure, scalable, highly available cloud infrastructure across GCP and AWS, supporting private, customer-managed, dedicated, and hybrid deployment models.
- - Design networking, IAM, compute, storage, database, container, and security architectures across development, QA, staging, production, and customer environments.
- - Establish cloud architecture standards, reusable patterns, and operational guardrails; review designs for reliability, security, scalability, performance, and cost.
- - Deploy and operate Kubernetes platforms using GKE and Amazon EKS, managing cluster configuration, node pools, autoscaling, ingress, networking, storage, secrets, certificates, and workload identity.
- - Package and deploy services using Docker and Helm; implement secure workload isolation and environment‑specific configuration.
- - Troubleshoot Kubernetes networking, scheduling, resource, storage, and deployment issues; improve cluster resilience, availability, performance, and cost efficiency.
- - Build and maintain reusable Infrastructure-as-Code modules using Terraform to automate cloud environments, networking, Kubernetes clusters, databases, storage, IAM, monitoring, and security controls.
- - Implement proper state management, module versioning, code review, testing, and promotion practices to prevent unmanaged changes and environment drift.
- - Maintain complete infrastructure documentation and deployment runbooks.
- - Design and maintain secure CI/CD pipelines for application, database, infrastructure, and configuration releases, automating build, test, security scanning, packaging, deployment, verification, and rollback.
- - Support release promotion across development, QA, staging, production, and customer environments, integrating database migrations, reference‑data changes, and configuration validation.
- - Establish immutable release‑artifact and environment‑promotion practices to improve release frequency, predictability, traceability, and recovery.
- - Apply security controls throughout cloud infrastructure and software‑delivery pipelines, including least‑privilege IAM, workload identity, secrets management, encryption, network controls, vulnerability scanning, and audit logging.
- - Support private VPC deployments, VPNs, private endpoints, firewalls, load balancers, DNS, and secure customer connectivity.
- - Automate container, dependency, infrastructure, and configuration security scanning; support enterprise and government security, audit, and compliance requirements.
- - Build observability using cloud‑native and open‑source logging, monitoring, tracing, and alerting tools; define SLIs, SLOs, dashboards, and actionable alerts.
- - Monitor infrastructure, Kubernetes, applications, databases, AI services, and deployment pipelines; lead investigation and resolution of complex production issues.
- - Conduct root‑cause analysis and implement…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).