Systems Automation Engineer
Listed on 2026-07-14
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
systems automation engineer
location: us-co-colorado springs
estimated starting salary range: usd $/yr.
- usd $/yr. Salary to be determined by the education, experience, knowledge, skills, and abilities of the applicant, internal equity, and alignment with market data.
the systems automation engineer architects, secures, and sustains a fully air-gapped, multi-account kubernetes platform in aws govcloud, built on rhel, rke2, and self-managed gitlab. This role is built on an automation-first mindset infrastructure-as-code, one‑shot repeatable deployments, ci/cd pipelines, and containerization to deliver highly scalable, resilient, and secure solutions in a disconnected environment. The engineer eliminates manual processes in favor of bundled, offline‑reproducible deployments, embeds devsecops and zero‑trust identity across the full system lifecycle, and operates the platform to meet nist, rmf, stig, and il5/il6/ts‑sci requirements.
platform& infrastructure engineering
- designing, deploying, and sustaining rhel 8/9/10 environments and rke2 kubernetes clusters across multiple aws accounts and vpcs
- operating the cluster control plane and per-environment agent pools (dev/test/staging/prod) with node labeling, taints, and workload isolation
- performing hardening, patching, and performance tuning in a disconnected environment using bundled rpm delivery (no upstream repos), selinux, and host firewalls
- administering platform-native core services coredns, aws route 53, ingress/load balancing, and certificate lifecycle (cert‑manager / tls secrets)
- designing ha/dr and resilience multi-node control planes, automated backup/restore of stateful workloads (pvcs, gitlab, databases), and failover across availability zones
- architecting and managing aws govcloud environments (ec2, s3, ebs, vpc, iam, route 53) across multiple accounts with cross‑account roles, key pairs, and ami strategy
- designing secure, scalable multi‑vpc network architectures transit gateway meshes, subnets, routing, and cidr‑based security‑group rules (cross‑vpc sg references are not available in govcloud)
- implementing all infrastructure as code with terraform (reusable per‑vpc modules, cross‑account providers, programmatically rendered variables)
- standing up and operating an in‑environment oci/container registry and bootstrap services for offline image distribution
- optimizing cloud cost, performance, and availability through right‑sizing, repeatable destroy/redeploy, and monitoring
- building and sustaining ci/cd pipelines on self‑managed gitlab ee, including version‑ladder upgrades, the container registry, and auto‑registered kubernetes runners
- integrating devsecops container/image cve scanning (e.g., trivy, grype, clair), secrets management, and policy/gating in pipelines
- operating container build and delivery tooling (docker/podman, kaniko, helm) with helm‑template‑derived, air‑gap‑reproducible image bundles
- implementing gitops delivery with argocd (pull‑based, per‑app applications, environment promotion and prod gating)
- developing automation that eliminates manual intervention single‑command, phase‑driven deployments validated end‑to‑end before delivery
- engineering self‑contained, offline deployment bundles (rpms, oci images, helm charts, binaries) for reproducible installs in disconnected networks
- writing and maintaining automation in ansible, python, and bash; automating provisioning, configuration management, and patching
- maintaining idempotent, resumable deploy/upgrade/teardown workflows
- implementing zero‑trust identity and sso with keycloak (oidc/saml), federating platform services (gitlab, argocd, monitoring, admin uis) and supporting cac/piv x.509 smart‑card authentication
- implementing system hardening in accordance with disa stigs and nist baselines as automated, repeatable controls
- supporting rmf/ato processes authoring control evidence, security documentation, poa&m items, and continuous‑monitoring inputs
- implementing centralized logging, monitoring, and alerting (elastic/siem, host/audit telemetry, and aws‑native sources where…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).