Sr Site Reliability Engineer
Job in
Westbrook, Cumberland County, Maine, 04092, USA
Listed on 2026-08-05
Listing for:
Artech
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Senior Site Reliability Engineer
Senior Site Reliability Engineer to focus on the health of a cloud-native event-driven enterprise transactional system supporting a billion-dollar line of business — its reliability, observability, performance, and resilience. As an embedded member of a large feature development team, this role brings dedicated, proactive attention to keeping the system healthy and operable, so that quality and stability advance in step with new features rather than trailing behind them.
RequiredSkills & Qualifications
- 7 years in an SRE, Dev Ops, infrastructure, or software engineering role.
- Experience tuning application performance based on real-world inputs — memory/CPU, message queues, request load — and adjusting resources and scaling accordingly.
- Proven, hands-on load and performance testing experience — designing realistic scenarios, executing load/stress tests, and interpreting results to drive capacity and performance decisions (Gatling preferred; comparable tools such as JMeter, k6, or Locust welcome).
- Hands-on Kubernetes experience.
- Strong Terraform / infrastructure-as-code experience.
- Solid working knowledge of Data Dog — comfortable enough to be the team's go-to and coach others on it.
- Comfortable defining and implementing SLOs and tuning alerting.
- Strong troubleshooting across the stack (we run Spring Boot services and an Angular frontend).
- Demonstrated success defining SRE and operational practices and driving their adoption across a development team.
- Genuine interest in applying AI to engineering work.
- Prior work experience at client or in client's Industry
- Champion and deepen our use of Data Dog across Logging, Error Tracking, RUM, Incident Management, and Case Management, and coach team members to raise their fluency with it.
- Improve signal-to-noise on alerts and errors; define and implement SLOs; make detecting and triaging issues faster.
- Own infrastructure-as-code: author and maintain Terraform and configure our Kubernetes service operator to manage and deploy infrastructure.
- Analyze infrastructure and application performance; tune sizing and scaling for cost-vs-performance, including refactoring application code where it improves reliability or performance.
- Own and evolve our Gatling load-testing framework: design realistic, high-volume load and stress scenarios, run them regularly, and translate the results into capacity, scaling, and performance decisions.
- Establish and lead our chaos engineering practice from the ground up — design and run fault-injection experiments and game days to validate resilience and systematically harden the system.
- Look for opportunities to embed AI/Claude into tooling and workflows to reduce manual toil.
- Monitor infrastructure cost trends and drive efficiency improvements.
- Participate in the on-call rotation: acknowledge alerts, triage, and coordinate the right people to remediate.
- Act as a leader during incident response — coordinating the response, ensuring stakeholders are kept informed, and pulling in the right people to resolve incidents within our recovery time objectives (RTOs).
- Triage production alerts and Tier-3 escalations during business hours; route issues and engage team members as appropriate.
- Document defects well in Jira (clear repro steps, recordings where useful) and initiate/coordinate post-mortems for critical incidents.
- Participate fully in team ceremonies (planning, grooming, retros).
- Help shape the SRE backlog — bringing your experience to bear on what we prioritize.
- Help define observability and supportability requirements for new features.
- Champion reliability and resilience practices and bring the rest of the team along.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×