Weekend Site Reliability Engineer
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-08-03
Listing for:
Sporty Group
Full Time
position Listed on 2026-08-03
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
What You’ll Be Doing
- Work with a team of Dev Ops and DBA professionals; covering Saturday, Sunday and Monday (5 days in total with flexibility in your days off) as a Weekend SRE
- Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future
- Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through Git Ops-first practices
- Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM)
- Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews
- Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding
- Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation
- Take ownership and responsibility for our cloud operation activities
- Liaise with external security agencies for annual audits as well as perform our own internal security sweeps
- Aid in reconfiguring existing architecture to allow for rapid deployments to new countries
- Mentoring less experienced team members
- 3+ years Dev Ops / platform engineering experience
- Must be based in Europe or Asia or LatAM
- Experience independently leading the planning and deployment of a project
- Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production
- Strong understanding of Kubernetes and container orchestration, with experience in EKS and Git Ops tooling such as ArgoCD and Helm being highly valued
- Experience with Infrastructure-as-Code, particularly Terraform
- Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus
- Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and Open Telemetry
- Experience with real user monitoring (RUM), with familiarity in Grafana Faro or Open Telemetry SDK instrumentation being a plus
- Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions
- Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns
- Experience defining SLIs and SLOs and using them to inform reliability work
- Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments
- Solid networking knowledge, especially the TCP / IP stack and HTTP protocol
- Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments
- A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached
- Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous
- Languages:
Java / Spring Boot, Node.js, Python, Java Script - Database:
Aurora MySQL & PostgreSQL, MongoDB, MySQL Community - Cache:
Elasti Cache, Redis, Valkey - Messaging:
Apache RocketMQ, AutoMQ, Kafka - Networking & Proxy:
Nginx, Kong, Cilium, eBPF - Orchestration & Git Ops:
Docker, Kubernetes (EKS), ArgoCD, Helm - Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3
- CI/CD:
Jenkins, Git Hub Actions - Metrics:
Prometheus, Mimir, Grafana, Alert manager - Logs:
Loki, Vector - Traces:
Tempo, Open Telemetry, Alloy - Profiling:
Pyroscope - RUM:
Grafana Faro, Open Telemetry SDK - Infrastructure as Code:
Terraform - CDN & Edge:
Cloudflare, AWS Cloud Front - AWS Cloud Watch
- Sporty is a remote first company in pursuit of sustainability
- A competitive salary + individual performance based bonuses every quarter
- 28 days paid annual leave
- Our core working hours are 10am-3pm in your local time zone with flexibility outside of this
- Referral bonuses & flash bonuses
- Top of the line equipment
- Annual company retreats to provide great internal networking opportunities
- Remote video screening with our Talent Acquisition Team
- Online assessment via Hackerrank
- Remote video interview with 3 x Team Members (45 mins each, not separate days)
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×