More jobs:
Observability Engineer
Job in
Raleigh, Wake County, North Carolina, 27601, USA
Listed on 2026-08-21
Listing for:
TEKsystems c/o Allegis Group
Full Time
position Listed on 2026-08-21
Job specializations:
-
IT/Tech
SRE/Site Reliability
Job Description & How to Apply Below
- Experience designing and implementing enterprise observability solutions across metrics, logs, and distributed traces.
- Experience developing instrumentation standards using Open Telemetry and modern observability frameworks.
- Experience building dashboards, alerts, and service health indicators that support application reliability and business outcomes.
Description
Our client is seeking a Senior Observability Engineer to help drive the evolution of enterprise observability capabilities across a large-scale cloud environment. This individual contributor will play a key role in developing observability standards, improving platform visibility, and enabling engineering teams to proactively monitor, troubleshoot, and optimize critical applications and services.
This position is ideal for someone who is passionate about monitoring, telemetry, distributed systems, and reliability engineering and enjoys partnering with development, SRE, and platform teams to improve operational excellence at scale.
Key Responsibilities:
Observability Engineering
Design and implement enterprise observability solutions across metrics, logs, and distributed traces.
Develop instrumentation standards using Open Telemetry and modern observability frameworks.
Build dashboards, alerts, and service health indicators that support application reliability and business outcomes.
Improve visibility across development, test, and production environments.
Partner with engineering teams to identify and close monitoring gaps.
Platform Optimization:
Support ongoing consolidation and modernization of observability tooling.
Develop strategies to improve signal quality while reducing telemetry noise.
Help establish governance standards for monitoring, alerting, and data ingestion.
Drive best practices around observability adoption across multiple engineering organizations.
Recommend improvements to platform reliability, scalability, and performance.
Reliability & Operations:
Support incident investigation through effective monitoring and root-cause analysis.
Develop service-level indicators (SLIs), service-level objectives (SLOs), and reliability metrics.
Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced observability practices.
Participate in architectural discussions around resiliency and operational readiness.
Cost Optimization:
Analyze telemetry consumption and platform usage trends.
Implement strategies for reducing unnecessary ingestion costs.
Help standardize logging, retention, sampling, and observability governance practices.
Balance visibility requirements with overall platform efficiency.
Required Qualifications
5+ years of experience in Observability Engineering, SRE, Dev Ops, Platform Engineering, or related disciplines.
Strong knowledge of:
Metrics
Logs
Distributed Tracing
Alerting & Monitoring Strategies
Hands-on experience with Open Telemetry and Prometheus.
Experience supporting cloud-native environments (AWS preferred).
Experience with observability platforms such as:
Coralogix
Datadog
Splunk
Grafana
New Relic
Understanding of microservices architectures and cloud-based systems.
Experience troubleshooting production environments using telemetry data.
Strong communication and collaboration skills.
Preferred Qualifications
Experience with Terraform or Infrastructure as Code.
Exposure to CI/CD platforms including Git Hub, Jenkins, or Azure Dev Ops.
Knowledge of SRE principles and operational excellence frameworks.
Experience supporting enterprise-scale observability initiatives.
Familiarity with platform engineering and developer enablement practices.
What Success Looks Like
Improved observability coverage across critical systems.
Reduced monitoring gaps between non-production and production environments.
Lower telemetry costs while maintaining platform visibility.
Faster incident detection and resolution.
Increased adoption of observability standards across engineering teams.
Skills
Linux, Cloud, Python, Aws, Devops, Automation, Azure, Administration, Active directory, Terraform, Kubernetes, Security
Top Skills Details
Linux,Cloud,Python,Aws,Devops,Automation,Azure,Administration,Active directory,Terraform,Kubernetes,Security
Additional Skills & Qualifications
N/A
Experience Level
Expert Level
Job Type & Location
This is a Contract to Hire position based out of Raleigh, NC.
Pay and Benefits
The pay range for this position is $70.00 - $85.00/hr.
Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors.
Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to specific elections, plan, or program terms. If eligible, the benefits available for this temporary role may include the following:
Medical, dental & vision Critical Illness, Accident, and Hospital…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×