×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer - Incident Insights & Readiness

Job in Boston, Suffolk County, Massachusetts, 02298, USA
Listing for: United States Digital Space LLC
Full Time position
Listed on 2026-10-08
Job specializations:
  • Software Development
    Cloud Engineer - Software, DevOps, Software Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way

The Incident Insights & Readiness SRE team at the company fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are the company engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across the company to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning.

Our efforts empower the company to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

* At the company, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.*

What You'll Do:
  • Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.
  • Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across the company. Our aim is to fully support our incident responders in dealing with complexity.
  • Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.
  • Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.
  • Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.
  • Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.
  • Lead cross-functional initiatives in engineering organizations across the company, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.
Who You Are :
  • At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of Type Script.
  • Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.
  • Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.
  • Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.
  • Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.
  • Empathy, collaboration, and communication…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary