×
Register Here to Apply for Jobs or Post Jobs. X

Observability Engineer

Job in Oklahoma City, Oklahoma County, Oklahoma, 73116, USA
Listing for: Oklahoma AG
Full Time, Part Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 85000 - 120000 USD Yearly USD 85000.00 120000.00 YEAR
Job Description & How to Apply Below
## Observability Engineer Apply locations:
Oklahoma County time type:
Full time posted on:
Posted Todaytime left to apply:
End Date:
August 17, 2026 (13 days left to apply) job requisition :
JR63293
** Job Posting Title
** Observability Engineer
** Agency
* * 090 OFFICE OF MANAGEMENT AND ENTERPRISE SERV
** Supervisory Organization
** IS-CS
** Job Posting End Date
** Refer to the date listed at the top of this posting, if available. Continuous if date is blank.

Note:

Applications will be accepted until 11:59 PM on the day prior to the posting end date above.
** Estimated Appointment End Date (Continuous if Blank)
**** Full/Part-Time
** Full time
** Job Type
** Regular
* * Compensation
* * Job Details  
• This is a full-time, 40-hour per week position  
• Support the Information Services Division  
• Position is on-site in Oklahoma City, OK
** Job Description
**** Job Details
*** This is a full-time, 40-hour per week position
* Support the Information Services Division
* Position is on-site in Oklahoma City, OK
** Position Summary
** The Observability Engineer is responsible for building, maintaining, and continuously improving the monitoring and observability capabilities that keep the organization's server, cloud, and network environments healthy. Using Datadog or similar platforms, this role instruments infrastructure and applications, builds meaningful dashboards and alerts, and turns raw telemetry into early warning signals that let teams find and fix problems before they impact the business.

This is a hands-on technical role for someone who enjoys making complex environments visible, understandable, and measurably more reliable.
** Position Responsibilities
**** Monitoring & Observability Engineering
*** Design, deploy, and maintain monitoring and observability tooling (Datadog or similar) across server, cloud, and network environments.
* Instrument infrastructure, applications, and services with metrics, logs, traces, and synthetic checks to provide full-stack visibility.
* Build and maintain dashboards that give clear, role-appropriate visibility into system health for engineers, managers, and leadership.
* Configure and tune alerting thresholds and escalation policies to catch real issues early while minimizing noise and alert fatigue.
* Integrate monitoring tools with incident management, ticketing, and on-call notification systems (e.g., Pager Duty, Service Now, Slack).
** IT Health & Continuous Improvement
*** Track and report on key IT health indicators such as uptime, latency, error rates, capacity headroom, and patch/compliance status across multiple environments.
* Partner with infrastructure, cloud, and network teams to identify recurring issues and drive root-cause fixes rather than repeated firefighting.
* Support capacity planning by analyzing utilization trends and flagging environments approaching risk thresholds.
* Contribute to post-incident reviews by providing telemetry, timelines, and health data that clarify what happened and why.
* Continuously refine monitoring coverage as new systems, services, and cloud resources are added, retiring stale checks and dashboards.
** Collaboration & Documentation
*** Work closely with server, cloud, network, and application teams to understand what 'healthy' looks like for each environment and translate that into monitoring coverage.
* Document monitoring standards, runbooks, and dashboard conventions so coverage stays consistent as the environment grows.
* Train and support other engineers in interpreting dashboards, alerts, and observability data.
* Evaluate and recommend improvements or additions to the observability toolset as monitoring needs evolve.
** Physical Demands and Work Environment
*** Office-based work involving extensive computer and phone use.
* Requires long periods of sitting, up to eight hours a day.
* Possible on-call rotation
* Work environment is generally quiet, occasional travel may be required.
** Education and Experience
*** Associate's or Bachelor's degree in Information Technology, Computer Science, or related field, or equivalent hands-on experience.
* 3+ years of experience in infrastructure monitoring, observability, systems administration, or IT operations.
* Hands-on experience with Datadog or a comparable observability platform (e.g., Dynatrace, New Relic, Splunk, Prometheus/Grafana).
* Working knowledge of server, cloud (AWS, Azure, or GCP), and network fundamentals sufficient to instrument and troubleshoot across environments.
* Experience building dashboards, alerts, and notification workflows that support fast, accurate incident response.
* Comfortable working with scripting or query languages (e.g., Python, Power Shell, Bash, DQL/PromQL) to build and refine monitoring logic.
** Preferred Qualifications
*** Datadog certification or equivalent vendor certification.
* Experience with infrastructure-as-code (Terraform, Ansible) for deploying monitoring agents and configuration at scale.
* Familiarity with ITSM practices (incident, problem, change management) and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary