Senior Observability Engineer
Listed on 2026-06-19
-
Software Development
Cloud Engineer - Software, DevOps
Job Summary
Math Works has a hybrid work model that enables staff members to split their time between office and home. The hybrid model provides the advantage of having both in-person time with colleagues and flexible at-home life optimizations. Learn More:
As a Senior Observability Engineer, you wil contribute to definig and driving the strategy, architecture, and implementation of observability capabilities across Math Works’ cloud platform and product engineering teams. This role is ideal for someone who brings deep expertise in cloud‑native observability systems, understands how modern distributed systems behave, and can lead cross-functional initiatives that improve reliability, performance, and end‑to‑end visibility for Math Works' online products and services.
Math Works nurtures growth, appreciates inclusivity, encourages initiative, values teamwork, shares success, and rewards excellence.
Responsibilities- Architect and evolve observability systems using cloud‑native tools such as Prometheus, Thanos, Alert manager, Open Telemetry
- Build scalable, multi‑tenant observability solutions for Kubernetes clusters running microservices at scale.
- Implement SLOs, SLIs, and error budgets—integrating observability into SRE practices.
- Improve signal quality and coverage by driving instrumentation standards across teams.
- Develop automation and tooling to enhance telemetry collection, alerting, and correlation.
- Support application teams in adopting observability patterns, instrumentation frameworks, and dashboards.
- Collaborate with incident management and SRE teams to improve root‑cause analysis and reduce MTTR.
- A bachelor's degree and 6 years of professional work experience (or a master's degree and 3 years of professional work experience, or equivalent experience) is required.
- Deep expertise in cloud-native observability stacks (Prometheus ecosystem, Open Telemetry, Grafana suite).
- Strong understanding of Kubernetes internals and common cloud‑native patterns (sidecars, operators, CRDs).
- Experience applying SRE principles: SLO/SLI design, chaos engineering, error budget management.
- Experience in designing self‑service observability capabilities for platform users.
- Security‑aware approach to telemetry data handling.
- Knowledge of cloud providers (AWS, Azure, GCP).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).