×
Register Here to Apply for Jobs or Post Jobs. X

AWS DevOps Datadog

Job in Atlanta, Fulton County, Georgia, 30383, USA
Listing for: NTT DATA North America
Full Time position
Listed on 2026-07-31
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 170000 USD Yearly USD 120000.00 170000.00 YEAR
Job Description & How to Apply Below

Company Overview

NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization,

NTT DATA's Client is currently seeking an AWS Dev Ops / SRE Engineer
- Datadog & AIOps

Location

Location:

Atlanta, Georgia
- Preferred Onsite/Hybrid

Employment Type

Employment Type:

Full-Time

Experience Level

Experience Level: 5+ years in Dev Ops or Cloud Engineering with production SRE experience

Role Summary

We are seeking a hands-on AWS Dev Ops Engineer with strong Site Reliability Engineering capabilities and deep Datadog experience. This role will design and improve secure, scalable CI/CD pipelines; increase platform reliability through observability, automation, and SLO-driven practices; and introduce practical AIOps and generative AI capabilities that improve build quality, deployment safety, incident response, and engineering productivity.

Day-to-Day

Job Duties
  • Design, build, and maintain resilient AWS environments using services such as EKS, EC2, S3, IAM, Lambda, RDS, Cloud Watch, Route 53, ALB/NLB, and Secrets Manager.
  • Build, standardize, and optimize CI/CD pipelines using Git Lab CI, Git Hub Actions, Jenkins, or similar platforms, with automated testing, quality gates, approvals, rollback, and progressive-delivery controls.
  • Apply SRE practices by defining service-level indicators, service-level objectives, error budgets, availability targets, and operational-readiness criteria.
  • Implement and administer Datadog capabilities including infrastructure monitoring, APM, log management, Real User Monitoring, synthetics, dashboards, monitors, service maps, and incident workflows.
  • Create actionable observability and alerting strategies that reduce noise, improve mean time to detect and recover, and support rapid root-cause analysis.
  • Automate infrastructure provisioning and configuration using Terraform, Cloud Formation, Ansible, or equivalent Infrastructure as Code tools.
  • Operate containerized workloads using Docker and Kubernetes/EKS, including autoscaling, health checks, resource optimization, and cluster reliability.
  • Integrate security and compliance controls into CI/CD, including secrets management, IAM least privilege, vulnerability scanning, SAST/DAST, dependency checks, artifact integrity, and audit evidence.
  • Use AIOps and generative AI to improve pipeline efficiency through intelligent failure analysis, configuration review, test generation, anomaly detection, change-risk scoring, and remediation recommendations.
  • Develop automation and operational tooling using Python, Bash, Power Shell, or similar scripting languages.
  • Lead production troubleshooting, incident response, post-incident reviews, problem management, and permanent corrective-action tracking.
  • Partner with application, platform, security, QA, and product teams to improve deployment frequency, change-failure rate lead time, reliability, and recovery performance.
  • Maintain runbooks, architecture diagrams, operational procedures, and engineering standards in Confluence, Jira, or similar tools.
  • Provide technical guidance and mentor engineers on cloud reliability, observability, automation, Dev Ops, and SRE practices.
Basic Qualifications
  • Minimum 5+ years of experience in Dev Ops, Cloud Engineering, Platform Engineering, or a related role supporting enterprise production systems.
  • Minimum 3+ years of hands-on experience designing, deploying, and operating solutions on AWS.
  • Minimum 5+ years of Strong experience building and supporting production-grade CI/CD pipelines;
    Git Lab CI experience is preferred.
  • Minimum 5+ years of Demonstrated SRE experience with SLOs/SLIs, error budgets, incident response, on-call operations, reliability engineering, capacity planning, and blameless post-incident reviews.
  • Minimum 5+ years of Strong hands-on Datadog experience across metrics, logs, APM/tracing, dashboards, alerting, monitors, synthetics, integrations, and service-level reporting.
  • Minimum 5+ years of Experience with Terraform or Cloud Formation and repeatable Infrastructure as Code practices.
  • Minimum 5+ years of Experience with Docker and Kubernetes; AWS EKS experience is strongly preferred.
  • Minimum 5+ years of Proficiency in at least one scripting or programming language such as Python, Bash, Power Shell, Go, or JavaScript/Type Script.
  • Strong understanding of Linux, networking, IAM, secrets management, cloud security, and production troubleshooting.
  • Minimum 5+ years of Experience working in Agile environments and using tools such as Jira, Confluence, and Git.
  • Strong communication, collaboration, documentation, and problem-solving skills with a high degree of ownership.
Preferred Qualifications
  • Hands-on experience applying AIOps, machine learning, or generative AI to CI/CD, observability, incident management, automated testing, code review, or root-cause analysis.
  • Experience integrating LLM-based assistants or agents with developer platforms, repositories, ticketing systems, observability…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary