Senior Staff DevOps Engineer — Data Discovery & AI Governance
Listed on 2026-08-08
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, Data Engineering
Strength in Trust
One Trust’s mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn’t slow teams down—it should accelerate what’s possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, One Trust is once again redefining what responsible innovation looks like.
One Trust, the AI‑Ready Governance Platform™, unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse sted by thousands of organizations worldwide, One Trust is shaping the future where trusted data becomes a transformative force for business and society.
One Trust's mission is to enable innovation through the responsible use of data and AI. We believe that ensuring data is trusted shouldn't slow teams down—it should accelerate what's possible. This led us to develop the first technology platform for responsible data use in 2016. Today, with AI representing the latest and most impactful expansion of data yet, One Trust is once again redefining what responsible innovation looks like.
One Trust, the AI-Ready Governance Platform™ unifies regulatory intelligence, automation, and connected governance workflows so businesses can continue to move at the speed of AI while ensuring good governance to prevent data misuse sted by thousands of organizations worldwide, One Trust is shaping the future where trusted data becomes a transformative force for business and society.
We are hiring a Senior Staff Dev Ops Engineer to join our Detect & Discover (D&D) team. This team owns three product lines — Data Discovery, Privacy Automation, and AI Governance — serving thousands of enterprise customers across multi‑cloud and on‑premises environments.
In this role, you will lead the infrastructure strategy for our Kubernetes‑based on‑premises platform and cloud deployments. This includes architecting automated worker node deployments for Azure Marketplace, AWS Marketplace, and GCP; driving container‑hardening initiatives to systematically eliminate CVEs; and ensuring the reliability, scalability, and security of our distributed scanning and classification platform. You will act as a technical leader, mentoring engineers and collaborating closely with product security, engineering squads, and customer‑facing teams.
YourMission
- Kubernetes Platform Architecture:
Architect and maintain production‑grade Kubernetes platforms across cloud (AKS, EKS, GKE) and on‑premises (K3s, microk8s) environments. Own the full lifecycle — cluster provisioning, upgrades, node pool management, and disaster recovery. - Deployment Automation & Marketplace Publishing:
Develop and maintain automation using Helm, ARM templates, BICEP, Terraform, and Cloud Formation to streamline worker node provisioning on Azure Marketplace, AWS Marketplace, and Reduce customer setup complexity so clients don't need advanced Kubernetes expertise on‑site. - Container Hardening & Vulnerability Management:
Lead our container‑hardening program by adopting secure base images for distributed components including Kafka, PostgreSQL, Elasticsearch, Temporal and Vault. Own the systematic reduction of CVEs flagged by Vulnerability Scanners. - CI/CD Releases Engineer:
Maintain and improve CI/CD pipelines ensuring zero‑downtime deployments, automated rollbacks, and full auditability. Drive Git Ops practices, Improve Dev Ex and infrastructure‑as‑code across the team. - Observability & Platform Reliability:
Design monitoring, alerting, diagnostics, and self‑healing capabilities using Datadog, Loki, Promtail. Ensure the platform is optimized for performance, logarithmic cost growth, and high‑throughput scaling across thousands of tenant environments. - Incident Response and Product Support:
Lead incident investigations, root cause analysis, and long‑term remediation for production service interruptions. Participate in on‑call rotation and drive systemic fixes to prevent recurrence. - Mentorship &…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).