Application Support DevOps Engineer
Listed on 2026-07-23
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Administrator, IT Infrastructure
With 1,000+ intelligence professionals serving over 1,900 clients worldwide, Recorded Future is the world’s most advanced, and largest, intelligence company!
Recorded Future equips our teams and customers with secure, isolated environments for high-risk research, threat analysis, and sensitive investigative work. To scale this capability, we rely on a container streaming platform that delivers browser-accessible desktops, applications, and isolated browsers on demand. As our operational footprint grows, we need an Application Support Dev Ops Engineer to own the deployment, tuning, and day‑to‑day health of this platform across our Kubernetes environments.
Your work will directly enable analysts and internal teams to spin up disposable, policy‑controlled work spaces, ensuring they have fast, reliable, and secure access to the tools they need without compromising our security posture.
- Lead the deployment, upgrade, and lifecycle management of a containerized workspace streaming platform running on Kubernetes, including Helm chart customization, namespace and zone design, and configuration of core services such as the API, manager, and connection proxy components.
- Design, build, and maintain custom Docker images (desktops, browsers, and single‑application work spaces) based on upstream core images, including adding software, startup scripts, branding, and launch forms to meet internal team requirements.
- Build and maintain custom VM images and auto‑scaling agent pools (on Kube Virt, Harvester, or cloud hypervisors) to support containerized sessions, RDP‑based sessions, and GPU‑accelerated workloads.
- Operate and scale the underlying infrastructure, including Kubernetes clusters, Linux servers, and cloud resources (AWS, Azure, GCP), covering capacity planning, node management, storage, and patching of agent hosts.
- Own networking and ingress for the platform: TLS certificate management (including cert‑manager integration), ingress controllers, load balancers, DNS, firewall rules, and secure routing between control plane, agents, and end users.
- Build CI/CD pipelines that automate image builds, vulnerability scanning, registry publishing, Helm chart deployments, and configuration rollouts across environments.
- Monitor platform health, investigate incidents, tune performance (session density, resource limits, autoscaling thresholds), and drive root‑cause analysis on deployment, session, and connectivity issues.
- Partner with Security, IT, and engineering teams to integrate the platform with SAML and OIDC identity providers, session recording, logging, and SIEM pipelines, and compliance controls.
- Document runbooks, deployment standards, and image‑build procedures, and mentor other engineers and support staff on operating the platform.
- 5+ years of experience in a Dev Ops, SRE, platform engineering, or application support role supporting production workloads.
- Strong hands‑on experience with Kubernetes (deployments, services, ingress, Helm, RBAC, troubleshooting pods and networking) in real production environments. Experience with Rancher, RKE2, EKS, AKS, or GKE is a plus.
- Deep Docker expertise, including authoring multi‑stage Dockerfiles, building and optimizing images, managing private registries, and debugging container runtime issues.
- Solid Linux systems administration skills (Ubuntu or Debian preferred), including scripting in Bash and at least one of Python or Go for automation and tooling.
- Working knowledge of at least one major cloud provider (AWS, Azure, or GCP), including compute, networking (VPC, subnets, security groups), storage, IAM, and managed Kubernetes offerings.
- Strong networking fundamentals, including TLS and PKI, DNS, HTTP and HTTPS proxies, reverse proxies, load balancers, and common troubleshooting tools (tcpdump, curl, dig, kubectl exec).
- Experience building CI/CD pipelines (Git Hub Actions, Git Lab CI, Jenkins, or similar) for image builds and Kubernetes deployments.
- Comfort operating and patching Linux VMs and servers in a production setting, including creating reproducible VM images (Packer, cloud‑init, or equivalent).
- Bonus:
Experience with Infrastructure‑as‑Code…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).