Senior Site Reliability Engineer
Listed on 2026-08-22
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
ISO New England Inc., One Sullivan Road, Holyoke, Massachusetts, United States of America
Job DescriptionPosted Monday, August 17, 2026 at 4:00 AM
ISO New England is the independent system operator responsible for ensuring the safe and reliable flow of electricity in our region and planning for the future of the electric grid. We are at the forefront of New England’s ongoing transition to clean energy.
The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions.
This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and Power Shell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.
What we offer you:
- A stable, mission-driven workplace where your impact truly matters
- A highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeing
- Competitive compensation with a base salary + performance bonus
- Robust benefits package, including:
- Enhanced 401(k) and financial planning support
- Tuition reimbursement and professional development
- Wellness programs, including an onsite gym
- Employee Business Networks
- Free coffee at our onsite café
- Hybrid work environment (3 days/week onsite)
- Distance-based relocation assistance available
How you will make an Impact
- Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, Ops Genie, Status Page)
- Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows
- Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead
- Develop meaningful KPIs and dashboards for business and IT service health
- Engineer and implement resilience patterns including HA, DR, and automated failover
- Partner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage
- Conduct performance testing, capacity modeling, forecasting, and right-sizing
- Participate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements
- Identify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions
- Design, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, Power Shell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency
- Design, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability
- Identify gaps in observability coverage and drive engineering solutions to close them
- Collaborate with architecture and application teams to ensure production readiness
- Leverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness
- Reduce repeat incidents by engineering permanent fixes and driving continuous improvement
What we are looking for
- 5+ years of experience in SRE, Dev Ops, systems engineering, platform engineering, or IT operations
- Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation.
- Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices
- Strong scripting and automation experience using Python, Power Shell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments
- Demonstrated experience designing,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).