Software Reliability Engineer
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer
Join The Un-carrier Movement
At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give our customers. All team members receive a competitive base salary and compensation package - this is Total Rewards. Employees enjoy multiple wealth-building opportunities through our annual stock grant, employee stock purchase plan, 401(k), and access to free, year-round money coaches. That's how we're UNSTOPPABLE for our employees!
Job Overview:
This role improves and protects software and systems supporting IT services by managing scalability, availability, latency, performance, security, and capacity. This role supports the Subscription Product Engineering organization, including in-house subscription and customer lifecycle platforms that support critical business operations and customer-facing services across production and non-production environments. The role primarily involves designing and maintaining continuous integration and continuous delivery, CI/CD, pipelines and building applications on cloud-native platforms.
The role differentiates itself by enabling continuous improvement of operational support through automation, monitoring, and reliability-focused practices across production and non-production environments. Success is measured by enhanced software delivery speed, reliability, operational efficiency, platform stability, and a consistent customer experience. We are a team that encourages innovation and advocates an agile and open approach, truly working and playing in the Un-carrier way!
Job Responsibilities:
- Apply Dev Ops automation tools to manage CI/CD pipelines and configuration for production and non-production environments.
- Perform environment management and automated server provisioning to support scalable infrastructure.
- Deliver software improvements that improve availability, scalability, latency, and efficiency of IT services.
- Create and manage dashboards, alerts, logging standards, and health checks to improve service quality, supportability, and visibility across services.
- Contribute to software delivery process improvements including cloud enablement, containerization, and deployment automation.
- Support cloud-native applications, APIs, microservices, and platform operations across production and non-production environments.
- Troubleshoot production incidents, participate in root cause analysis, and support implementation of long-term reliability improvements with assistance from leadership and senior technical team members.
- Partner with Software Engineering, Dev Ops, and platform teams to improve application resiliency, scalability, and deployment automation under established technical direction.
- Contribute to operational readiness activities, including release validation, capacity planning, disaster recovery support, and environment support, under the guidance of senior leadership.
- Participate in Agile ceremonies, production support activities, and continuous improvement initiatives.
- Also responsible for other duties/projects as assigned by business management as needed.
Education and
Work Experience:
- Bachelor's Degree plus 2 years of related work experience
OR combination of education and experience deemed equivalent (Required) - 2-4 years Relevant experience. (Preferred)
- Experience working in an Agile and Dev Ops environment. (Preferred)
- Experience in one or more of: C, C#, Java, Perl, Python, Go, or scripting experience in Shell and Perl. (Preferred)
- Experience in Continuous Integration/Continuous Delivery tools, such as, Jenkins, Cloudbees, etc., and other automation tools. (Preferred)
- Experience with Dev Ops tools, such as, Ansible, Chef, Puppet, etc. Experience in Docker, Kubernetes, etc. is preferable. (Preferred)
- Experience in APM tool, like, App Dynamics, logging tool, like Splunk. (Preferred)
- Experience working in a cloud environment (public/private). (Preferred)
- Experience in migrating to cloud or cloud native environments. (Preferred)
- Experience supporting APIs, microservices, distributed applications, or enterprise production platforms. (Preferred)
- Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).