Site Reliability Engineer/Irving , TX; Onsite
Listed on 2026-07-24
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Site Reliability Engineer
Location:
Irving, TX (Onsite)
*** 2 Positions available
Hands-on experience in deploying and managing Kubernetes clusters in telecom industry is required.
Experience with a cloud-native CI/CD tool used for Kubernetes deployments is necessary.
Knowledge of GKE (Google Cloud Platform) and RKE (Rancher Kubernetes Engine) is preferred.
Knowledge of Rook-Ceph distributed storage is preferred.
You should be able to manage multi-region clusters for disaster recovery.
You should have experience working with programmable infrastructure, such as building a CI/CD pipeline.
Experience with monitoring and observability tools such as Prometheus, Grafana, and the ELK stack is expected.
You should possess good knowledge of Linux operating systems and be proficient in troubleshooting OS issues.
It is essential not to use terms like high availability or resilient systems without a thorough understanding of their basics, as building such systems in practice requires significant effort.
Knowledge of security and compliance frameworks and best practices relevant to financial institutions is necessary.
6+ years of professional experience as SRE
Experience in Site Reliability Engineering.
Splunk/App Dynamic else other Application monitoring tool sets
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).