×
Register Here to Apply for Jobs or Post Jobs. X

Operational Data & Observability Engineer

Job in Seattle, King County, Washington, 98127, USA
Listing for: Nscale
Full Time position
Listed on 2026-07-17
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 145000 - 180000 USD Yearly USD 145000.00 180000.00 YEAR
Job Description & How to Apply Below

Operational Data & Observability Engineer

US

About the Role

We're looking for an Operational Data & Observability Engineer to build and evolve the monitoring, logging, and observability capabilities that power our production environments. In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights.

You’ll partner closely with Dev Ops, Site Reliability Engineering (SRE), platform, and software engineering teams to develop modern observability solutions, improve incident response, and enable data‑driven operational excellence.

What You’ll Do Design & Build Observability Solutions
  • Design and implement enterprise observability strategies across infrastructure, services, and applications.
  • Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility.
  • Build and maintain centralized logging and log analysis pipelines.
  • Implement distributed tracing to improve visibility across microservices and complex application workflows.
  • Establish performance baselines and develop anomaly detection strategies.
Operational Data Engineering
  • Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems.
  • Design and manage operational data pipelines that support monitoring and analytics.
  • Develop APIs and integrations that enable operational data consumption across teams.
  • Ensure data quality, consistency, retention, and cost‑efficient storage practices.
Reliability & Operations
  • Troubleshoot production issues using monitoring, logging, and tracing data.
  • Participate in an on‑call rotation and support incident response activities.
  • Create and maintain operational documentation, runbooks, and troubleshooting guides.
  • Partner with engineering teams to improve platform reliability, scalability, and operational readiness.
  • Continuously optimize observability infrastructure for performance and resilience.
Platform & Tool Administration
  • Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies.
  • Evaluate emerging observability tools and recommend improvements.
  • Automate monitoring deployments, instrumentation, and platform configuration.
  • Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure.
What You’ll Bring

Required Qualifications
  • 3+ years of experience in Dev Ops, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering.
  • Hands‑on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent.
  • Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, Cloud Watch, or similar solutions.
  • Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent.
  • Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM).
  • Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies.
  • Solid understanding of application, infrastructure, networking, database, and storage performance monitoring.
  • Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem‑solving.
Preferred Qualifications
  • Experience supporting microservices‑based architectures.
  • Expertise across multiple observability platforms.
  • Experience with incident management, root cause analysis, and post‑incident reviews.
  • Infrastructure as Code experience using Terraform, Ansible, or similar tools.
  • Familiarity with eBPF or low‑level Linux performance monitoring.
  • Experience building custom telemetry, ETL, or operational data pipelines.
  • Understanding of security monitoring, audit logging, and compliance requirements.
What Success Looks Like
  • Improve platform visibility and operational health.
  • Reduce Mean Time to Resolution (MTTR) during incidents.
  • Increase alert quality while reducing unnecessary noise.
  • Deliver highly available, scalable observability…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary