Site Reliability Engineer; SRE AI & Automation
Job in
Dallas, Dallas County, Texas, 75201, USA
Listed on 2026-09-02
Listing for:
Iconma
Full Time
position Listed on 2026-09-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Site Reliability Engineer (SRE) – AI & Automation
Our client, an IT Services and Consultant company, is looking for a Site Reliability Engineer (SRE) – AI & Automation for their Dallas, TX / Scottsdale, AZ / Hybrid location.
Responsibilities:
- Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
- Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
- Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
- Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
- Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
- Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and App Dynamics to meet reliability and performance objectives.
- Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Requirements:
- Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
- Kubernetes Platform Engineering – 5+ years of strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
- Cloud & Infrastructure Automation – Strong experience in GCP, Terraform, Helm, Git Hub, CI/CD, and production-grade automation.
- Software Development – 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
- Observability & Monitoring – Splunk, Grafana, Datadog, App Dynamics, alerting, and platform health monitoring.
- API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
- AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
- Years of
Experience:
14.00 Years of Experience
Why Should You Apply?
- Health Benefits
- Referral Program
- Excellent growth and advancement opportunities
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×