Intermediate Observability Engineer
Listed on 2026-09-06
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
At CI, we see a great place to work as one that is a safe place for everyone to have a voice, where people are empowered to take ownership over meaningful work, where there is an opportunity to grow through stretching themselves, where they can work on innovative products and projects, and where employees are supported and engaged in doing so.
We are seeking a Mid-Level Observability Engineer to help build, maintain, and enhance our enterprise monitoring and observability capabilities across cloud and hybrid environments. This role is hands‑on and execution‑focused, supporting Dynatrace and AWS Cloud Watch implementations, dashboard development, alert tuning, instrumentation, and operational reporting for critical platforms and applications
The ideal candidate will partner with Cloud Engineering, Dev Ops, SRE, and application teams to improve service visibility, strengthen monitoring coverage, and embed observability practices into ongoing operational and delivery workflows.
Key Responsibilities- 1. Monitoring & Platform Operations Assist in implementing and maintaining centralized observability capabilities using Dynatrace and AWS Cloud Watch across enterprise systems and cloud workloads. Deploy, configure, and support Dynatrace One Agent and related monitoring components across hybrid infrastructure and modern application environments. Help onboard applications, services, and infrastructure into monitoring platforms, ensuring basic visibility for availability, performance, logs, metrics, and traces.
Support the creation and maintenance of dashboards, management zones, tagging structures, notebooks, and alert configurations. - 2. APM, Alerting & Troubleshooting Configure and maintain monitoring for distributed applications, APIs, databases, and cloud‑native workloads. Build and tune persona‑based dashboards and alerting thresholds for engineering, operations, and leadership audiences while helping reduce unnecessary alert noise. Support incident triage and root cause analysis using Dynatrace features such as Davis AI, topology mapping, and distributed tracing under the guidance of senior engineers.
Assist with post‑incident follow‑up by identifying monitoring gaps and recommending dashboard, threshold, or instrumentation improvements. - 3. Cloud & Telemetry Engineering Implement cloud observability capabilities using AWS Cloud Watch, including logs, metrics, alarms, dashboards, and basic analytics. Support distributed tracing and log aggregation for microservices, containerized platforms, and serverless workloads such as EKS, ECS, and Lambda. Work with development and Dev Ops teams to improve instrumentation consistency and embed monitoring tasks into CI/CD workflows. Contribute to automation of monitoring deployment and configuration tasks through scripting and infrastructure‑as‑code practices.
- 4. Operational Excellence & Team Collaboration Create recurring operational reports and dashboards that track application health, infrastructure performance, and service‑level indicators. Participate in monitoring reviews, tuning exercises, and platform housekeeping activities to improve quality of observability data Collaborate with senior engineers and cross‑functional teams to document standards, improve adoption, and support enterprise monitoring initiatives.
3–5 years of experience in Observability, Monitoring, Dev Ops, Systems Engineering, or SRE‑related roles.
Hands‑on experience with Dynatrace, including deployment, configuration, dashboarding, alerting, and basic APM support.
Working knowledge of AWS infrastructure and monitoring concepts, especially Cloud Watch and common cloud services used in enterprise environments.
Experience with scripting or automation using Python, Bash, or Power Shell for monitoring operations and platform support tasks.
Familiarity with Terraform or Cloud Formation for basic infrastructure and configuration deployment.
Understanding of observability fundamentals, including metrics, logs, traces, alerting, the golden signals, and structured logging practices.
Good understanding of modern application environments such as microservices, container platforms, and distributed…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: