Senior DevOps Engineer
Listed on 2026-09-05
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
We are looking for talented Platform/Dev Ops engineer with deep expertise in Kubernetes, Database technologies, Terraform, and Ansible to help scale ••••• ’s AI platform across on-premises, cloud, and SaaS environments.
The right candidate will design infrastructure and deploy ••••
• software to support our US Government customers in fully realizing the value of AI.
This role demands a strong foundation in Linux, networking (both traditional and Kubernetes), container technologies, and automation. You’ll collaborate closely with internal and external engineering teams, own critical infrastructure, and solve challenging operational and scalability problems in fast-paced, dynamic environments.
From your first day, you will make a valuable — and valued — contribution. We are a fast-growing company where no one is a bystander. We offer you the opportunity to delight millions of consumers around the world while gaining meaningful experience across a variety of disciplines.
We are looking for talented Platform/Dev Ops engineer with deep expertise in Kubernetes, Database technologies, Terraform, and Ansible to help scale ••••• ’s AI platform across on-premises, cloud, and SaaS environments.
The right candidate will design infrastructure and deploy ••••
• software to support our US Government customers in fully realizing the value of AI.
This role demands a strong foundation in Linux, networking (both traditional and Kubernetes), container technologies, and automation. You’ll collaborate closely with internal and external engineering teams, own critical infrastructure, and solve challenging operational and scalability problems in fast-paced, dynamic environments.
From your first day, you will make a valuable — and valued — contribution. We are a fast-growing company where no one is a bystander. We offer you the opportunity to delight millions of consumers around the world while gaining meaningful experience across a variety of disciplines.
Duties and Responsibilities:
- Design, architect, and operate highly available, multi-tenant Kubernetes platforms across cloud and on-premises environments.
- Own the full networking stack — CNI, service mesh, ingress, DNS, load balancing, and network policy.
- Own database operations for production Postgres and Open search clusters — availability, performance, backup, and recovery.
- Partner with engineering teams to define platform and delivery standards.
- Automate infrastructure provisioning, configuration, and lifecycle management.
- Enforce security best practices across the platform — network policies, RBAC, secrets management, and vulnerability patching.
- Lead incident response and post-mortems for platform-level failures; implement systemic fixes.
- Proactively identify and remediate scalability and reliability risks across the platform.
Skills and
Qualifications:
- Active U.S. DoD Top Secret clearance required
- 5+ years in Platform Engineering, Infrastructure Engineering, or Dev Ops supporting large-scale distributed systems.
- 5+ years of Kubernetes experience — cluster architecture, multi-tenancy, RBAC, scheduling, security standards, and autoscaling across cloud and bare-metal. Experience with k3s is a plus.
- Strong Kubernetes networking knowledge — CNI (Calico, Cilium), service mesh (Istio, Traefik), ingress controllers, and Network Policy.
- Linux networking fundamentals — TCP/IP, DNS, BGP, and network troubleshooting.
- Experience designing and operating infrastructure across hybrid environments — on-premises, edge, and multiple cloud providers (AWS, Azure, OCI).
- Infrastructure as code proficiency — Terraform and Ansible.
- Working knowledge of Helm/Kustomize for application packaging and deployment.
- Proficiency in Python, Go, and Bash for automation and tooling.
- Postgres experience — replication, HA/failover, connection pooling (PgBouncer), query tuning, backup/recovery, and Kubernetes operators (CloudNative PG).
- Experience operating stateful workloads on Kubernetes including Elasticsearch/Open search and Postgres.
- Experience operating storage solutions (CSI, Rook/Ceph).
- Cloud-native observability experience — Prometheus, Grafana, Loki, and Tempo.
- Experience with Security Standards — FIPS, CVE mitigation,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).