×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer SRE Weekend Coverage

Job in Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: Framework Ventures
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Site Reliability Engineer (SRE) – Weekend Coverage

About Elwood:
We have built a digital asset trading infrastructure for institutional investors. Our seamless end‑to‑end platform connects to global crypto exchanges, custodians, and liquidity providers, via a single API. Built by industry experts with decades of combined experience in investment management and digital technology, Elwood provides market infrastructure at scale, enabling financial institutions, neobanks, and corporations to access digital asset markets quickly and efficiently.

We are seeking a Site Reliability Engineer (SRE) to join our globally distributed engineering team, with a key responsibility for weekend operations and system reliability. You’ll play a critical role in maintaining uptime, resolving incidents, and automating infrastructure for Elwood’s EMS and PMS platforms, which are built on AWS and GCP cloud environments. This highly visible role blends deep technical ownership with cross‑functional collaboration.

In addition to core SRE responsibilities, you will support our Technical Account Managers and client‑facing teams in resolving production issues that impact users, ensuring smooth and reliable client experiences. You’ll be part of a team responsible for both the infrastructure backbone and the performance reputation of Elwood’s platform.

Key Responsibilities
  • Ensure the reliability, availability, and performance of production systems, particularly during weekends.
  • Take ownership of monitoring, troubleshooting, and incident response during weekends and off‑hours.
  • Actively participate in on‑call rotations with priority weekend shifts (Saturday–Sunday).
  • Troubleshoot and resolve critical issues in a fast‑paced, high‑availability environment.
  • Automate manual processes and workflows, reducing operational overhead.
  • Work closely with engineering teams to design and deploy scalable, fault‑tolerant infrastructure solutions on AWS or GCP.
  • Improve observability by utilizing monitoring, logging, and alerting systems (e.g., Cloud Watch, Datadog).
  • Lead post‑incident reviews, contribute to the continuous improvement of system reliability, and follow up on strategic fixes.
  • Develop and update runbooks, incident response playbooks, and documentation.
  • Work closely with Engineering, Product, and Client teams to proactively identify infrastructure pain points that could affect the user experience.
  • Monitor alert channels, logs, and infrastructure load for the entire stack.
  • Set up automation for alerting.
Required Experience
  • 5+ years of experience in an SRE, Dev Ops, or infrastructure engineering role.
  • Strong experience with AWS or GCP, including services like EC2, Lambda, S3, RDS, and GKE (for GCP).
  • Experience with automation tools like Terraform.
  • Proficient in at least one scripting language (Python, Bash, Go, etc.).
  • Solid understanding of Linux systems, networking, and cloud‑based architectures.
  • Experience working with container orchestration platforms like Kubernetes.
  • Proficient with CI/CD pipelines, preferably with cloud‑native tools (e.g., Git Hub).
  • Ability to troubleshoot complex, distributed systems and provide solutions in high‑pressure environments.
  • Ability to communicate effectively with both technical and non‑technical stakeholders.
Preferred Qualifications
  • Exposure to Execution Management Systems (EMS) / Portfolio Management Systems (PMS).
  • Experience with client‑impact triage, working cross‑functionally with account managers or product teams.
  • Proficiency with Datadog or similar observability platforms.
  • Knowledge of serverless architectures (e.g., AWS Lambda, GCP Cloud Functions).
  • Familiarity with RDBMS and No

    SQL databases, such as RDS, CloudSQL, DynamoDB.
  • Prior experience in fintech, trading platforms, or 24/7 financial infrastructure.
  • Strong understanding of API integrations and how infrastructure issues might manifest in client environments.
  • Excellent problem‑solving and communication skills, with the ability to translate technical incidents into clear client updates.
  • Experience working with client‑facing teams.
Working Hours and Availability

The successful candidate must be comfortable working autonomously during weekend shifts (Thursday‑Monday…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary