Senior Observability Engineer
Listed on 2026-07-10
-
Software Development
DevOps, Software Architect, Cloud Engineer - Software, Backend Developer
We are looking for a Senior Observability Engineer to define and implement enterprise-wide telemetry architecture standards for cloud-native and AI-driven platforms. The role focuses on Open Telemetry, ADOT, AWS observability services, distributed tracing, AI/LLM observability, and cross-platform telemetry governance. The ideal candidate combines deep hands-on engineering expertise with the ability to establish observability standards across multiple teams and platforms.
Project OverviewOur customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting‑edge technology. Our teams are involved in the engineering of core platform components for best‑in‑class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
Requirements- Open Telemetry architecture (SDK, Collector, semantic conventions — including GenAI SIG conventions for AI workloads)
- AWS Distro for Open Telemetry (ADOT) — collector pipeline design, processor configuration, multi-exporter setup
- AWS Cloud Watch (Logs Insights, EMF, composite alarms, cross-account aggregation)
- AWS X‑Ray (distributed tracing, service maps, trace groups, sampling rules)
- New Relic (APM, distributed tracing, custom instrumentation, NRQL)
- Open Telemetry GenAI SIG semantic conventions
- Agent/LLM observability (Lang Smith, MLflow tracing, or equivalent)
- Cross‑runtime observability architecture (multi‑vendor telemetry consolidation patterns)
- Python (ADOT instrumentation, SDK development, integration with agent frameworks)
- 5+ years of observability or platform engineering with hands‑on Open Telemetry implementation
- AWS Cloud Watch and X‑Ray at scale (multi‑account, cross‑service correlation)
- New Relic in enterprise environments (APM, custom dashboards, alerting)
- Open Telemetry GenAI SIG semantic conventions or AI/LLM observability implementation
- Agent or LLM system observability (tracing LLM calls, tool invocations, token usage)
- Technical leadership translating enterprise requirements into actionable engineering plans
- Cross‑team coordination on observability standards across multiple engineering teams
- AWS Agent Core Observability
- Experience with Datadog or Dynatrace in hybrid environments
- Strands or Lang Graph agent framework instrumentation
- Design and maintain enterprise observability architecture based on Open Telemetry standards.
- Define telemetry collection, enrichment, processing, and export strategies across cloud-native and AI workloads.
- Build and optimize ADOT and Open Telemetry Collector pipelines, including processors, exporters, sampling strategies, and multi‑destination telemetry routing.
- Establish observability standards, semantic conventions, and instrumentation guidelines across engineering teams.
- Design distributed tracing solutions using AWS X‑Ray, Open Telemetry, and New Relic.
- Implement and maintain enterprise observability capabilities in AWS Cloud Watch, including Logs Insights, EMF, composite alarms, and cross‑account observability.
- Develop telemetry instrumentation and integrations using Python and Open Telemetry SDKs.
- Define and implement observability standards for AI applications, agent frameworks, LLM workloads, tool invocations, and model interactions.
- Implement tracing and monitoring for agent-based systems, including token consumption, LLM inference calls, workflow execution, and tool execution visibility.
- Create cross‑runtime observability strategies spanning multiple telemetry vendors and execution environments.
- Partner with platform, security, AI, and engineering teams to drive telemetry governance and operational excellence.
- Build dashboards, alerting strategies, troubleshooting workflows, and observability best practices for enterprise‑scale environments.
- Lead observability architecture reviews and translate business requirements into actionable implementation plans.
- Mentor engineering teams on instrumentation standards and telemetry adoption.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).