×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Remote / Online - Candidates ideally in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listing for: HostPapa Inc.
Remote/Work from Home position
Listed on 2026-07-21
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 100000 - 130000 USD Yearly USD 100000.00 130000.00 YEAR
Job Description & How to Apply Below

Position Summary

With team members and customers in 39 countries around the globe, Host Papa is currently one of the fastest-growing web hosting companies with a wide range of products available. At its core, we provide individuals and small and medium-sized businesses with access to valuable tools and services critical to their online success, including a Website Builder service for making website creation an ultra-easy task for anyone.

Tailored to meet every user's unique needs, our award-winning customer support, email, and cloud-based solutions keep Host Papa at the cutting edge of the web hosting industry and innovation by putting our customers first.

This role focuses on Cloud Blue, a Host Papa business that powers cloud commerce for many of the world’s largest service providers, including major Telcos, distributors, and MSPs. Cloud Blue enables partners to monetize and manage cloud services and subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization.

As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of Cloud Blue’s multi-tenant SaaS platforms used by service providers worldwide. You will focus on improving system stability and performance through monitoring, high availability, and incident response, while working closely with Dev Ops, Platform, and Engineering teams to build and operate resilient production systems.

What you’ll do
  • Define and implement SLIs, SLOs, and error budgets for critical Cloud Blue services to ensure reliability and performance
  • Influence system architecture with a strong focus on reliability, scalability, and operability, designing systems for fault tolerance, graceful degradation, and self-healing
  • Reduce operational toil by identifying opportunities for automation and process improvement
  • Design and operate Cloud Blue’s observability stack across metrics, logs, and traces using tools such as Datadog, Grafana, and Elastic Stack
  • Develop actionable alerting strategies and dashboards that provide clear insight into platform and business health
  • Design and maintain high-availability architectures, implementing redundancy, failover, and disaster recovery strategies across regions and availability zones
  • Conduct capacity planning, load testing, and performance optimization to ensure platform stability and scalability
  • Act as a senior responder during production incidents, leading incident coordination, communication, and service restoration
  • Own blameless postmortems and drive improvements that reduce incident frequency, MTTR, and customer impact
  • Improve reliability of Kubernetes-based platforms through health checks, autoscaling strategies, rollout safety, and resilience testing
  • Partner with engineering and Dev Ops teams to improve deployment safety, rollback strategies, and platform reliability
  • Maintain runbooks and operational documentation, and promote SRE best practices across engineering teams
  • Support other tasks or projects as assigned to meet team and business needs
About you
  • 3+ years of experience as an SRE, Dev Ops Engineer, or Production Engineer, with strong ownership of production systems
  • Proven experience operating highly available, enterprise-grade, multi-tenant SaaS platforms
  • Hands-on experience with observability and monitoring tools such as Datadog, Grafana, and Elasticsearch/Kibana
  • Solid understanding of Linux, networking, and distributed systems fundamentals
  • Experience working with containerized environments such as Docker and Kubernetes
  • Strong scripting and automation skills using Python and/or Bash
  • Experience participating in on-call rotations and incident response in production environments
  • Strong written and spoken English
  • Experience defining SLIs/SLOs and managing error budgets at scale will be considered a plus
  • Exposure to hyperscale or service-provider-grade platforms is an advantage
  • Cloud experience, preferably with Azure; experience with AWS and/or GCP will also be valued
  • Experience working with hybrid or on-premises integrations is beneficial
  • Familiarity with chaos engineering and resilience testing will be considered an asset
What We Offer
  • Work from anywhere -…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary