Site Reliability Engineer
Job in
Irving, Dallas County, Texas, 75084, USA
Listed on 2026-07-27
Listing for:
Resolve Tech Solutions
Full Time
position Listed on 2026-07-27
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AWS, Systems Engineer
Job Description & How to Apply Below
Serve as a primary responder for production incidents, ensuring rapid triage, mitigation, and resolution.
Responsibilities- Lead root cause analysis (RCA) and drive long‑term corrective actions.
- Maintain and improve incident response processes, runbooks, and escalation paths.
- Collaborate with engineering, QA, and product teams to prevent recurrence of issues.
- Support and optimize AWS services such as EC2, ECS/EKS, Lambda, S3, Cloud Watch, IAM, RDS, and VPC networking.
- Monitor system health, performance, and capacity across cloud environments.
- Implement infrastructure best practices around reliability, scalability, and cost efficiency.
- Assist with deployments, environment configuration, and CI/CD pipelines.
- Manage and troubleshoot MongoDB clusters, including performance tuning, replication, backups, and failover.
- Diagnose query performance issues and collaborate with developers on schema optimization.
- Ensure data integrity, availability, and recovery readiness.
- Use New Relic, Cloud Watch, and other observability tools to monitor application and infrastructure performance.
- Build dashboards, alerts, and telemetry that provide actionable insights.
- Continuously refine monitoring thresholds to reduce noise and improve signal quality.
- Experience with on-call rotations and 24/7 production environments.
- Work cross-functionally with the various teams in the organization and help establish SLOs and achieve those SLOs.
- 5+ years of experience in SRE, Dev Ops, Production Support, or similar operational roles.
- Strong hands-on experience with AWS services and cloud-native architectures.
- Proficiency with MongoDB administration and troubleshooting.
- Experience with New Relic or similar APM/observability platforms.
- Experience using additional tools like Postman, Intune, and Firebase, Service Now, Cloudwatch.
- Strong understanding of Linux systems, networking, and distributed systems.
- Solid scripting skills (Python, Bash, or similar).
- 5+ years Monitoring and Alarming in all environments and familiar with tools like Mongo Charts, New Relic, Cloudwatch, Service Now.
- Proven experience managing high-severity incidents and driving RCA processes.
- Familiarity with CI/CD tools (Jenkins, Git Hub Actions, Git Lab CI, etc.).
- Experience with container orchestration (ECS, EKS, Kubernetes).
- Knowledge of message queues (Kafka, SQS, RabbitMQ).
- Exposure to microservices architectures.
- Certifications such as AWS Solutions Architect, AWS Sys Ops, or MongoDB DBA.
- Working experience with IoT devices, and Microsoft Intune.
[Pay range or salary or compensation]
Equal Opportunity Statement[Include a statement on commitment to diversity and inclusivity.]
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×