Sr Lead-Systems
Listed on 2026-09-02
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Role description
Job Title:
Systems Reliability Engineer
Location:
Texas, USA
Job Summary:
We are seeking a highly skilled and motivated Systems Reliability Engineer to join our dynamic team as a Technical Team Lead. The ideal candidate will possess extensive experience in Site Reliability Engineering (SRE) and will be responsible for ensuring the reliability, availability, and performance of our systems and applications. This role requires a strong technical background, particularly in Kubernetes, Docker, and Golang, along with proficiency in monitoring and automation tools such as Grafana, Bash, or Power Shell.
Responsibilities:
- Lead the SRE team in implementing best practices for system reliability and performance.
- Design, build, and maintain scalable and resilient infrastructure using Kubernetes and Docker.
- Develop and maintain automation scripts and tools to streamline operations and improve efficiency.
- Monitor system performance and reliability using Grafana and other monitoring tools, responding to incidents and outages as they occur.
- Collaborate with development teams to ensure that applications are designed with reliability and scalability in mind.
- Conduct post incident reviews and implement improvements to prevent future occurrences.
- Provide technical guidance and mentorship to junior team members.
- Stay current with industry trends and emerging technologies to continuously improve our systems.
Mandatory
Skills:
- Strong experience in Site Reliability Engineering (SRE) principles and practices.
- Proficiency in container orchestration using Kubernetes.
- Experience with Docker for containerization.
- Strong programming skills in Golang.
- Familiarity with monitoring and visualization tools, particularly Grafana.
- Proficient in scripting languages such as Bash or Power Shell.
- Excellent problem-solving skills and the ability to work under pressure.
- Strong communication and collaboration skills.
Preferred
Skills:
- Experience with cloud platforms (AWS, Azure, GCP).
- Knowledge of CI CD pipelines and Dev Ops practices.
- Familiarity with configuration management tools (e.g., Ansible, Puppet).
- Understanding of networking concepts and protocols.
- Experience with database management and optimization.
Years of
Experience:
- Must have at least 6 years of Experience
- Must have billing experience in Telecom Domain
Qualifications:
Bachelor's degree in Computer Science, Information Technology, or a related field. Relevant certifications in cloud technologies or SRE practices are a plus.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).