Site Reliability Engineer- Eng
Job in
Lowell, Middlesex County, Massachusetts, 01856, USA
Listed on 2026-07-13
Listing for:
Ukg-6
Full Time
position Listed on 2026-07-13
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Job description
Company and benefits
Job IDSTAFF
017752
Employment Type
Regular Work Stylehybrid Location Lowell ,MA,United States Travel Up to 25%
Role Staff Site Reliability Engineer
- Eng Why UKG:
At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our workforce operating platform. Helping people get paid, grow in their careers, and shape the future of their industries. That’s what we do.
We never stop learning. We never stop challenging the norm. We push for better, and we celebrate the wins along the way. Here, you’ll get flexibility that’s real, benefits you can count on, and a team that succeeds together. Because at UKG, your work matters—and so do you.
About the Team Staff Site Reliability Engineers (SREs) at UKG are senior individual contributors who play a critical role in ensuring the reliability, scalability, and performance of our services. They bring a breadth of knowledge across service delivery and apply software engineering principles to operational this role, you will ensure the reliability, availability, and performance of production systems by applying software engineering practices to operations.
SREs proactively monitor system health, manage risk through SLOs and error budgets, lead incident response, and enable safe, rapid change while balancing reliability and delivery velocity.
Staff SREs are passionate about learning and evolving with modern technologies. They strive to innovate and relentlessly pursue an excellent customer experience, with an “automate everything” mindset that enables services to be delivered with speed, consistency, and high availability.
This is a senior individual contributor role, focused on technical leadership, influence, and reliability impact.
About the Role and
Job Responsibilities Engage in and improve the lifecycle of services from conception to end-of-life, including system design reviews, capacity planning, and production readiness.
Define and implement standards and best practices for system architecture, service delivery, reliability, and automation, including the definition and monitoring of service health indicators (latency, traffic, error rates, and resource saturation), service level objectives (SLOs), and the use of error budgets to guide operational and delivery decisions.
Support service, product, and engineering teams by providing common tooling and frameworks to increase availability and improve incident detection and response.
Improve system performance, availability, and efficiency through automation, process refinement, post-incident reviews, and in-depth configuration analysis.
Collaborate closely with engineering teams across the organization to deliver and operate reliable services.
Increase operational efficiency, effectiveness, and service quality by treating operational challenges as software engineering problems (reducing toil).Guide junior team members and serve as a champion for Site Reliability Engineering best practices.
Actively participate in incident responses, including on-call rotations and post-incident reviews, collaborating with engineering teams to restore service and reduce recurrence.
Partner with stakeholders to influence and help drive the best possible technical and business outcomes.
Required Qualifications 5+ years of hands-on experience in software engineering, systems engineering, or cloud-based environments.
5+ years of experience working with public cloud platforms (e.g., GCP (preferred), AWS, or Azure).5+ years of experience configuring, operating, and maintaining applications and/or systems infrastructure in a large-scale, customer-facing environment.
Demonstrated understanding of observability best practices, including metric generation and collection, log aggregation pipelines, time-series databases, and distributed tracing.
Experience coding in one or more higher-level programming languages (e.g., Python, Java, or C++).Strong working knowledge of Linux systems, including troubleshooting, performance analysis, and scripting in production environments.
Experien…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×