Senior Reliability & Optimization Engineer
Job in
Englewood, Arapahoe County, Colorado, 80151, USA
Listed on 2026-10-05
Listing for:
Kforce Inc
Full Time
position Listed on 2026-10-05
Job specializations:
-
IT/Tech
SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Responsibilities
Kforce is immediately seeking an experienced Senior Reliability & Optimization Engineer in support of our enterprise telecommunications and mass media client based in Greenwood Village, CO. Responsibilities:
Platform Engineering (Primary):
- Maintain and extend the Terraform modules that define our infrastructure;
Audit AWS against state, detect drift, and reconcile it rather than letting the console diverge - Own the health of the Git Lab CI/CD pipelines and enforce pipeline-only deployment with progressive promotion (dev, qa, uat, stage, prod);
Never skip environments or patch production in the console - Operate the AWS footprint: EKS (blue/green upgrades, node groups, IRSA), Helm, and Istio;
Data and messaging (Aurora, Document DB, Redis, Amazon MQ);
Networking and edge (Route
53, WAFv2, Cloud Front);
And storage (S3) - Keep resources current and tagged to standard and remediate scan-flagged vulnerabilities through IaC
- Build, deploy, and validate releases across environments, including off-hours windows;
Document release notes and runbooks
Kforce is immediately seeking an experienced Senior Reliability & Optimization Engineer in support of our enterprise telecommunications and mass media client based in Greenwood Village, CO. Responsibilities:
Platform Engineering (Primary):
- Maintain and extend the Terraform modules that define our infrastructure;
Audit AWS against state, detect drift, and reconcile it rather than letting the console diverge - Own the health of the Git Lab CI/CD pipelines and enforce pipeline-only deployment with progressive promotion (dev, qa, uat, stage, prod);
Never skip environments or patch production in the console - Operate the AWS footprint: EKS (blue/green upgrades, node groups, IRSA), Helm, and Istio;
Data and messaging (Aurora, Document DB, Redis, Amazon MQ);
Networking and edge (Route
53, WAFv2, Cloud Front);
And storage (S3) - Keep resources current and tagged to standard and remediate scan-flagged vulnerabilities through IaC
- Build, deploy, and validate releases across environments, including off-hours windows;
Document release notes and runbooks
- Right-size and scale resources to meet SLAs at the lowest sustainable cost, distinguishing real demand growth from regressions and leaks, and recommending configuration changes that keep latency and spend on target
- Partner with developers and test engineers to improve application performance and strengthen automated test coverage for reliability-sensitive changes
- Own monitoring and alerting end to end: build Data Dog/Splunk dashboards, tune thresholds and routing against SLAs, and pre-empt degradation
- Provide first response, mitigation, recovery, and root-cause analysis for incidents and outages;
Participate in on-call rotation and drive follow-ups to closure
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience
- AWS operations:
Strong hands-on experience operating AWS through both the console and Infrastructure as Code, across services such as EKS, S3, Document DB, Aurora/RDS, Elasti Cache Redis, Lambda, Route
53, WAFv2, and Amazon MQ - CI/CD:
Experience maintaining Git Lab (or comparable) CI/CD pipelines for infrastructure and application deployments, including multi-stage environment promotion - Containerized microservices:
Production experience operating containerized microservice and web-based applications (Docker/Kubernetes) - Incident response:
Demonstrated experience owning production incident triage, mitigation, and root-cause analysis under SLA pressure - Performance work:
Experience benchmarking, performance testing, and optimizing resource usage and system behavior - Observability:
Expertise with monitoring and alerting tooling such as Data Dog and/or Splunk - building dashboards, tuning alert thresholds, and using telemetry to investigate latency, errors, and saturation - Source control:
Familiar with Git-based source control and branch/merge-request workflows (e.g., Git Lab) - Infrastructure as Code:
Proficient with Terraform - reading, maintaining, and authoring modules;
Managing state; and detecting and reconciling drift;
Comfortable making changes through code and pipelines rather than the console
The pay range is the lowest to highest compensation we reasonably in good faith believe we would pay at posting for this role. We may ultimately pay more or less than this range. Employee pay is based on factors like relevant education,…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×