Sr. Site Reliability Engineer
Listed on 2026-07-11
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AWS
Sr. Site Reliability Engineer
Length: 12+ Month contract
Location:
6380 S Fiddlers Green Cir, Greenwood Village, CO 80111
Onsite Requirements: while on contract – a few times a week until ramped up then can do once or twice a month
Project Scope/
Roadmap:
- Configuration Management Team is an internal experimentation and A/B testing platform that is implemented across all of Charter’s customer platforms and tools. This tool helps teams validate their releases so it doesn’t impact the customer experience negatively, and also ensures that enhancements being made to the products are truly moving the needle in a positive direction for the customer.
- This team also owns all the back end for TDCS (Targeted Delivery Client Services) Platform which is a real time configuration management tool used throughout all of Charter’s client facing platforms (streaming apps, customer web portals, etc.). All applications use this tool to deliver client specific experiences. The platform can target specific customers, markets, etc. to run experiments or conduct testing in order to control variants.
- This role specifically will be focused on supporting TDCS. They're moving everything to their own AWS account to help better manage their own releases and must maintain the availability and latency. This platform is crucial to all customers' ability to access streaming/video as it's called on at least twice a day from every application.
Must Haves:
- 6+ years of Dev Ops experience within large, complex environments. Ideally comes from a development background (able to read code and work alongside the developers).
- AWS and Terraform
- Kubernetes
- Gitlab or similar CI/CD tools for pipeline deployment
- Datadog or similar monitoring tools
Team/Environment:
- Team of 7 engineers and 1 lead
Job Description:
The applied AI and data science program at Charter brings together data scientists, data engineers, and software engineers in close collaboration to empower Spectrum employees to safely release and test product changes and understand the impact on customers. Join us to build technology that enables Charter’s product teams to deliver high-quality experiences across streaming video, broadband internet, and mobile services.
As a Senior SRE Engineer in the applied AI and data science program, you’ll help deploy, monitor, support, and optimize software solutions that expand Charter’s custom experimentation and configuration management platforms. You’ll work with cloud-based applications and resources hosted in AWS. The applications and services are written using Node.js, React, and Python, the infrastructure uses Terraform, and Git Lab is used to manage the deployment pipelines.
The CMX team is responsible for enhancing and supporting Charter’s custom, enterprise-level experimentation and configuration management platforms. We facilitate experimentation and configuration management through web-based UIs and supporting services that enable users to draft and deploy A/B tests, progressive deployments, and JSON configuration changes.
Responsibilities:
- Release Management
- Build and deploy application, service, and infrastructure releases
- Validate system operation and integrity post-deployment
- Document release notes
- Production Support
- Maintain 99.999% availability of critical production systems
- Ensure infrastructure and applications run smoothly
- Keep infrastructure resources updated
- On-call for production outage and incident response
- Perform root cause analysis for all production outages and incidents
- Monitoring and Alerting
- Apply monitoring and alerting policies to every system and provide recommendations
- Build and expand monitoring dashboards
- Monitor and log errors, bugs, and unexpectant behavior
- Monitor system efficiency including latency and resource consumption
- Monitor for system degradation, pre-empt or mitigate as possible
- Alert the appropriate people when monitoring detects operations outside of expected SLAs
- Optimization
- Manage resource scaling to align with project goals
- Optimize resource usage and system behavior
- Team Participation
- Assist with user support
- Become the resident expert of the system architecture, deployment pipelines, and resource utilization
- Coordinate efforts with onshore and offshore teammates
- Develop bug fixes
Primary
Qualifications:
- Expertise using monitoring tools such as Data Dog and/or Splunk
- Experience with AWS infrastructure and services (e.g. EKS, S3, Document DB)
- Experience working with containerized microservice and web-based applications
- Proficient with the AWS console and Infrastructure as Code (e.g. Terraform)
- Experience benchmarking and performance testing
- Experience deploying cloud-based applications
- Familiarity with Git-based source control and branch management (e.g. Git Lab)
- Bachelor’s degree in a related field or equivalent experience
Secondary
Qualifications:
- 6 or more years of professional experience with the software development lifecycle
- Familiarity with Python, Node.js, React, Type Script, and GraphQL
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).