×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer; SRE AI & Automation

Job in Dallas, Dallas County, Texas, 75201, USA
Listing for: Iconma
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE) – AI & Automation

Site Reliability Engineer (SRE) – AI & Automation

Our client, an IT Services and Consultant company, is looking for a Site Reliability Engineer (SRE) – AI & Automation for their Dallas, TX / Scottsdale, AZ / Hybrid location.

Responsibilities:

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and App Dynamics to meet reliability and performance objectives.
  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Requirements:

  • Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Kubernetes Platform Engineering – 5+ years of strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
  • Cloud & Infrastructure Automation – Strong experience in GCP, Terraform, Helm, Git Hub, CI/CD, and production-grade automation.
  • Software Development – 5+ years of advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
  • Observability & Monitoring – Splunk, Grafana, Datadog, App Dynamics, alerting, and platform health monitoring.
  • API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
  • AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
  • Years of

    Experience:

    14.00 Years of Experience

Why Should You Apply?

  • Health Benefits
  • Referral Program
  • Excellent growth and advancement opportunities
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary