More jobs:
Site Reliability Engineer (SRE
Job in
New York City, Richmond County, New York, USA
Listed on 2026-08-05
Listing for:
Georgia IT, Inc.
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Job Description & How to Apply Below
Site Reliability Engineer (SRE)
We are seeking an experienced Site Reliability Engineer (SRE) with strong expertise in Dynatrace to join our growing engineering team. The ideal candidate will be responsible for ensuring the reliability, scalability, performance, and observability of mission-critical applications and infrastructure across cloud and distributed environments. This role requires a hands-on engineer with deep experience in monitoring, automation, incident management, and cloud-native technologies.
Key Responsibilities- Design, implement, and manage end-to-end monitoring and observability solutions using Dynatrace.
- Configure dashboards, alerts, anomaly detection, and root cause analysis for complex distributed systems.
- Monitor application performance, infrastructure health, and system reliability across cloud environments.
- Collaborate with Dev Ops, Cloud, and Engineering teams to improve system availability and operational efficiency.
- Automate operational tasks and incident response workflows using scripting and Automation tools.
- Support and troubleshoot Linux/Unix servers, containers, and Kubernetes-based environments.
- Build and maintain CI/CD pipelines to streamline deployments and reduce operational overhead.
- Perform performance tuning, capacity planning, and reliability improvements.
- Implement best practices for observability, logging, monitoring, and incident management.
- Participate in on-call rotations and production support activities when required.
- Analyze production incidents and conduct root cause analysis (RCA) with preventive action plans.
Skills & Qualifications
- 6–10 years of experience in Site Reliability Engineering, Dev Ops, or Production Support roles.
- Strong hands-on expertise with Dynatrace, including:
- End-to-end monitoring
- Alerting and problem detection
- Dashboard creation
- Application Performance Monitoring (APM)
- Root cause analysis
- Solid understanding of observability, logging, and monitoring frameworks.
- Experience working with cloud platforms such as AWS, Azure, or GCP.
- Strong Linux/Unix administration and troubleshooting skills.
- Experience with containerization and orchestration technologies:
- Docker
- Kubernetes
- Proficiency with CI/CD tools such as:
- Jenkins
- Git Lab CI/CD
- Azure Dev Ops
- Strong scripting and automation skills using Python, Bash, or Shell scripting.
- Good understanding of microservices architecture and distributed systems.
- Experience with incident management, reliability engineering, and operational excellence practices.
- Strong analytical, troubleshooting, and communication skills.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×