More jobs:
DevOps Engineer
Job in
Roswell, Fulton County, Georgia, 30076, USA
Listed on 2026-07-18
Listing for:
Agile Resources, Inc.
Full Time
position Listed on 2026-07-18
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Job Description & How to Apply Below
Employment Type:
Permanent / Direct Hire.
Compensation:$120k–$130k (depending on experience).
No sponsorship will be offered for this role now or in the futureThis Senior Dev Ops Engineer role owns the reliability, automation, and day-to-day operability of a Kubernetes-based platform. Success requires deep, hands‑on cluster administration and troubleshooting, plus the ability to improve CI/CD and Infrastructure as Code practices with minimal direction. You will assess what’s in place, drive best‑practice improvements, and establish clear standards that enable engineering teams to deliver safely and consistently.
Responsibilities:- Take ownership of the Kubernetes platform by assessing the current cluster architecture, identifying gaps, and driving remediation through to stable, supportable outcomes.
- Lead cluster-level troubleshooting and root‑cause analysis across networking, storage, scheduling, resource constraints, and workload failures; implement preventative fixes to reduce repeat incidents.
- Build, maintain, and improve automated CI/CD pipelines in Git Hub to standardize build, test, and deployment workflows and shorten release cycles.
- Implement Infrastructure as Code to provision and manage platform and environment configurations in a repeatable, version‑controlled way.
- Establish and enforce Dev Ops standards and best practices, including pipeline conventions, deployment patterns, and operational guardrails; provide clear technical direction when tradeoffs arise.
- Strengthen Git‑based delivery practices by improving how configuration and deployment changes are reviewed, promoted, and audited across environments.
- Develop runbooks and operational documentation that make the platform easier to support, troubleshoot, and evolve over time.
- Partner with engineering teams to remove friction in build and deployment processes, resolve pipeline and platform issues, and improve delivery reliability.
- 5+ years of experience in Dev Ops, Site Reliability Engineering (SRE), Platform Engineering, or a similar role.
- Deep hands‑on administration of Kubernetes environments, including cluster-level troubleshooting/debugging, architecture understanding, and root‑cause analysis.
- Strong CI/CD experience, including building and maintaining automated build/test/deploy pipelines using Git Hub (e.g., Git Hub Actions/Workflows).
- Experience implementing Infrastructure as Code (IaC) (e.g., Terraform, Helm, or Ansible).
Skills:
- Experience operating Red Hat Open Shift clusters in production.
- Hands‑on AWS/EKS experience, including planning or supporting a cloud migration to AWS.
- Experience supporting Kubernetes platforms hosted on IBM Cloud.
- Experience implementing Git Ops workflows with ArgoCD.
- Experience with Flux or comparable Git Ops tooling for declarative delivery.
- Experience using Kustomize to manage environment overlays and configuration.
- Experience packaging and deploying applications using Helm charts.
- Experience leading or supporting Kubernetes/cloud migration efforts, including cutover planning and risk mitigation.
- Experience working in healthcare, financial services, or other regulated environments with strong audit and change‑control expectations.
- Strong Git Ops experience managing declarative infrastructure and application deployments end‑to‑end.
- Experience supporting Maven‑based build and dependency management workflows.
- Ability to develop automation/tooling in Python to reduce manual operational effort and improve platform reliability.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×