Site Reliability Engineer
Job in
Quincy, Norfolk County, Massachusetts, 02171, USA
Listed on 2026-09-02
Listing for:
TechDigital Group
Full Time
position Listed on 2026-09-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Mandatory
Skills:
Python/R and ML libraries (scikit-learn, Tensor Flow, PyTorch), Data analysis and visualization (Pandas, Num Py, Power BI/Tableau), SQL and database management
Key Responsibilities
- Monitor, maintain, and improve the reliability, availability, and performance of production systems.
- Design and implement monitoring, alerting, logging, and observability solutions.
- Establish and track Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Automate operational tasks and repetitive processes using scripting and Infrastructure as Code (IaC).
- Lead incident response activities, troubleshooting, root cause analysis (RCA), and post-incident reviews.
- Collaborate with development, infrastructure, and platform teams to improve system reliability and resilience.
- Perform capacity planning, performance tuning, and scalability assessments.
- Support CI/CD pipelines and deployment automation initiatives.
- Implement high-availability, disaster recovery, and failover strategies.
- Strong experience with Linux/Unix administration.
- Proficiency in scripting languages such as Python, Shell, or Power Shell.
- Hands‑on experience with cloud platforms (AWS, Azure, or GCP).
- Experience with containerization technologies such as Docker and Kubernetes.
- Knowledge of monitoring and observability tools such as Prometheus, Grafana, ELK, Splunk, Dynatrace, or Datadog.
- Understanding of CI/CD tools such as Jenkins, Git Hub Actions, Git Lab CI, or Azure Dev Ops.
- Experience with Infrastructure as Code tools such as Terraform, Ansible, or Cloud Formation.
- Strong troubleshooting, debugging, and problem-solving skills.
- Understanding of networking, security, and distributed systems concepts.
- 5–10+ years of overall IT experience.
- 5+ years of hands‑on experience in Site Reliability Engineering, Production Support, Dev Ops, or Cloud Operations roles.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×